24 KiB
Inline math rendering — framing
Status: framing only; no implementation.
This note frames a feature that currently does not exist: rendering
LaTeX math expressions as typed math (not raw source) in any buffer.
It proposes a four-tier pipeline and identifies the substrate changes
needed in the Rust core, the semantic protocol, and the GPU frontend.
The TUI frontend is explicitly scoped out (terminal cells cannot
express positioned math glyphs at interactive rates). If this is
adopted, the TUI path would be "display the raw $...$ source with a
distinct face as a lossy fallback" and nothing more.
The note assumes no dependency on KaTeX / MathJax / JavaScript runtimes. Everything from parsing through layout to rendering is native Rust.
Design contract (inherits the existing semantic protocol boundary)
From docs/semantic-frontend-protocol.md: the instance never learns a
pixel. Inline math does not reopen this. The instance detects math
ranges, parses LaTeX, and produces a position-independent math layout
tree; the frontend positions glyphs at pixel coordinates using its own
font metrics and viewport. The instance-to-frontend contract for math
is a list of positioned-glyph or font-size-scalable records, not a
pre-rasterized bitmap.
Contract additions specific to math
-
Math is text, not an image. A math expression is selectable, copy-pastes as its LaTeX source, and re-renders on edit with no round-trip. Invariant: the raw
$...$text in the rope is always the canonical source; the rendered glyphs are a projection. -
Cache invalidation is the frontend's job. The instance does not track which expressions are visible or dirty. The GPU frontend maintains a math layout cache keyed by source-text hash; only the expression under the cursor invalidates on each keystroke.
-
Display math is a block, not an inline layer.
$$...$$regions introduce vertical space and center the formula. They collapse to a single "block" cell in the base text layout and are rendered as a full-width overlay in a separate pass.
Four-tier architecture
buffer text
|
[1. Detection] — regex / tree-sitter → byte ranges tagged Math
|
[2. Parsing] — recursive-descent → MathNode tree
|
[3. Layout] — MATH-font metrics → positioned glyph runs
|
[4. Rendering] — wgpu / glyphon pass at computed coordinates
Tier 1: Detection
Every buffer, after every edit, scan for math delimiters. The default set:
| Delimiter | Kind | Notes |
|---|---|---|
$...$ |
inline | Single-dollar, non-greedy |
$$...$$ |
display | Double-dollar, greedy across newlines |
\(...\) |
inline | Alt inline (LaTeX convention) |
\[...\] |
display | Alt display |
The detection layer emits byte ranges with a tag:
enum MathKind { Inline, Display }
struct MathSpan { start: BytePos, end: BytePos, kind: MathKind }
Where it runs:
- For buffers with a tree-sitter grammar: a
(math_expression)node in the grammar signals scanned injection ranges (reuses the existingParseTreeBundle+Layermachinery fromdocs/multi-language-injections-framing.md). Each grammar that can contain LaTeX (markdown, org, raw TeX, etc.) needs a(math_expression) @mathcapture rule added — either by patching the upstream grammar or shipping an overlayqueries/math/highlights.scm. The framing defers enumerating which grammars need this until adoption; markdown and LaTeX grammars are the obvious v0 targets. - For buffers with no grammar: a fast byte-level scanner in Rust
(two-pass: find
$/$$/\(/\[boundaries, match pairs, handle escapes). Runs in the GPU frontend'srebuild_code_slice()(pmacs-gpu/src/main.rs:4849) as a post-shape hook, not the edit path, so keystroke latency is unaffected. (The TUI'ssrc/text_view.rsis a separate line-index view that does not participate.)
Cost note: the scan runs on every buffer after every edit,
regardless of whether the buffer is likely to contain LaTeX. A log file
with a bare $ will be scanned, find no pair, and exit. This is a
negligible cost per edit — the scan is O(length of changed region), not
O(file) — but the framing notes it as a minor inefficiency. An
extension-based gate (only scan buffers whose language is in a
configured set) is a trivial v1 optimization.
Extensibility: A Lua hook pmacs.math.delimiters registers
additional patterns per mode. Example:
pmacs.math.add_delimiter("markdown", "`$", "`$") -- `` $...$ `` in md
Tier 2: Parsing
A recursive-descent parser converts the LaTeX math source into an AST. LaTeX math mode is a constrained grammar — far smaller than full LaTeX. The node set covers what appears in real mathematical writing:
enum MathNode {
/// Characters and identifiers
Char(char),
Identifier(SmolStr), // \sin, \alpha, x
/// Subscript / superscript
Sub(Box<MathNode>), // _{...}
Super(Box<MathNode>), // ^{...}
SubSuper(Box<MathNode>, Box<MathNode>), // _{...}^{...}
/// Fractions
Fraction(Box<MathNode>, Box<MathNode>),
/// Radicals
Sqrt(Box<MathNode>),
SqrtN(Box<MathNode>, Box<MathNode>), // \sqrt[n]{...}
/// Big operators (sum, prod, int — limits above/below)
BigOp {
op: MathOp,
lower: Option<Box<MathNode>>,
upper: Option<Box<MathNode>>,
sub: Option<Box<MathNode>>, // \sum_{i=1} vs \sum_{i=1}^\infty
sup: Option<Box<MathNode>>,
},
/// Fences that may stretch
Fenced {
left: Delimiter,
body: Box<MathNode>,
right: Delimiter,
},
/// Accents
Accent { accent: AccentKind, base: Box<MathNode> },
/// Style overrides (e.g. \displaystyle)
Styled(Box<MathNode>, MathStyle),
/// \text{...} — literal text in math mode
Text(String),
/// Generic stretchy: \overline, \underbrace, etc.
Stretch(StretchKind, Box<MathNode>),
/// Sequences and groups
Group(Vec<MathNode>),
}
~500 lines of Rust. No lookup tables beyond symbol names → Unicode
codepoints (the \alpha → U+03B1 map is ~200 entries for Greek
- Hebrew + arrows + operators). The parser does not handle macro definitions, preamble material, or LaTeX3 — only math-mode markup.
Tier 3: Layout
This is the hard part and the place where pmacs would do something no interactive editor currently does natively: position math glyphs using an OpenType MATH table. The pipeline:
A. Font metrics. Load a math font with an OpenType MATH table. The candidates, ranked:
| Font | Table quality | License | Notes |
|---|---|---|---|
| Latin Modern Math | Full | OFL (GUST) | Reference, ships with TeX Live, most widely tested |
| STIX Two Math | Full | OFL | Broader Unicode coverage |
| Cambria Math | Full | Proprietary | Ships with Office; unavailable on Linux |
| Libertinus Math | Full | OFL | Derivative of Latin Modern, wider |
Bundled default: Latin Modern Math. The existing font-loading path
(build_font_system() → fontdb::Database::load_font_source() in
pmacs-gpu via cosmic-text) already handles .ttf/.otf; adding the MATH
table means pulling in one of ttf-parser or read-fonts from
fontations as a new dependency (neither is in the tree today — both are
pure Rust, low-risk additions).
B. Box model. Each MathNode lays out into a MathBox:
struct MathBox {
width: f32,
height: f32,
ascent: f32, // above baseline
descent: f32, // below baseline
italic_correction: f32,
glyphs: Vec<PosGlyph>,
}
struct PosGlyph {
glyph_id: GlyphId,
x: f32,
y: f32, // relative to box baseline
font_size: f32,
color: Color, // usually inherited from theme
}
Layout rules per node type (following Knuth's math layout algorithm, simplified to the common cases):
-
Group: lay out children left-to-right, accumulate width. Insert italic corrections between adjacent slanted glyphs (from the MATH table's
MathItalicsCorrection). -
Fraction: lay out numerator and denominator at 70% font size, centered horizontally; draw a rule between them at the math axis height (from MATH table
MathConstants::AxisHeight); the box ascent = num.ascent + axis + rule_thickness/2, descent = den.descent + space- rule_thickness/2.
-
Subscript / Superscript: scale to ~70%, shift superscript up by
SuperscriptShiftUp(from MATH table), shift subscript down bySubscriptShiftDown. If both present, adjust so they don't overlap. -
BigOp: select the display-size glyph variant from the MATH table (
GlyphVariantRecordchain). Place lower limit below, upper above, usingDisplayOperatorMinHeightfor minimum size. -
Fences (
\left(...\right)): measure the enclosed box height; select the smallest fully-enclosing glyph variant from the stretchy chain; if the height exceeds the largest single glyph, assemble from theGlyphConstructionparts (top, bottom, repeatable extender). -
Sqrt: lay out the radicand; draw the radical sign extending from the top-left to cover the radicand height, using the
RadicalKern,RadicalExtraAscender, andRadicalRuleThicknessconstants from the MATH table.
C. Caching. The key performance insight: a laid-out MathBox is
hashable (by the source LaTeX string). The cache is a HashMap<u64, Arc<MathBox>> keyed by SipHash of the source bytes. The GPU frontend
holds this cache across reshape calls. On edit, only the expression
whose source range overlaps the edit range is evicted.
Latency target: < 500 µs for the common case (a 20-node expression). Worst case (a full-page display equation with nested fractions and summations): < 5 ms. Cache hit: < 1 µs.
Bulk-doc note: for a LaTeX document with hundreds of inline expressions on screen simultaneously (e.g., a dense math paper at high zoom-out), each visible expression is parsed and laid out independently. The 5 ms worst case per expression could add up: 50 visible expressions at 1 ms each = 50 ms. The hash cache absorbs re-parses across frames (an expression re-appears on re-scroll at hash cost only), so the expensive path is only the first render of each expression after an edit or a fresh scroll. For v0 this is acceptable; if it proves hot, v1 can add a frame budget — parse until deadline, render cached results for the rest.
Tier 4: Rendering (GPU frontend)
The GPU frontend (pmacs-gpu/src/main.rs) already renders text through
cosmic-text + glyphon with per-span styling. Math expressions are a new
layer inserted into the existing z-order:
Backgrounds → Squiggles → Code Text → Math → Gutter → Caret → Minimap → Minibuffer → Completion → Context menu
Inline math ($...$):
rebuild_code_slice()detectsMathSpans in the visible range.- For each inline span, extract the source text, hash-lookup the math cache, parse+layout on miss.
- Delete the raw
$...$glyphs from the cosmic-text buffer. Insert a zero-width placeholder with explicit width =MathBox.width. - Collect all
PosGlyphruns into aVec<MathDraw>list, offset by the line's baseline position. - In
render(), after the code text draw call, iterateMathDraws and issue glyphonTextAreacalls for each glyph run at its computed absolute position.
Display math ($$...$$):
- The math source occupies full lines. Detect via the text layout that
the
MathSpanspans entire visual lines. - Replace the affected visual lines with a single spacer glyph.
- Insert vertical space before and after (from
\abovedisplayskipand\belowdisplayskipequivalents — hardcoded constants are fine for v0). - Render the math block centered horizontally at the spacer position.
Stretchy delimiters (the hardest rendering case): the MATH table's
GlyphConstruction entries describe how to assemble a vertically
stretched glyph from top/middle/bottom/extender pieces. The render pass
for a stretchy glyph draws 3–5 separate glyph IDs at computed
positions, bottom-to-top. This is analogous to the existing squiggle
shader (a custom WGSL path for diagnostic underlines); a stretchy-glyph
shader or vertex-buffer assembly follows the same pattern.
Protocol surface
No new wire types for v0. The GPU frontend detects math locally from
the text content it already receives via BufferSnapshot. The semantic
protocol stays unchanged; math rendering is a pure frontend
responsibility in v0.
Note: this means frontend-local detection duplicates regex scanning for every reshape. In the common case (a handful of visible math spans) the cost is negligible; for a full-screen display of a LaTeX document with hundreds of inline expressions the scan cost accumulates linearly with visible byte count. The hash cache absorbs re-parse of unchanged expressions, but the delimiter scan itself is unavoidable. If this proves hot, v1 moves detection to the instance side.
If this proves out, v1 would add an optional MathSpans variant to
InstanceMessage so the instance's tree-sitter detection (tier 1) is
authoritative and the frontend does not reimplement detection. This is
deferred until there is a frontend consumer to validate the wire shape.
As a secondary benefit, a single instance-side scan serves all connected
frontends — the v0 approach pays the scan cost per frontend.
Integration points
| Component | Change | Risk |
|---|---|---|
src/math_parse.rs (new) |
~500-line recursive-descent parser | Low; pure fn, no deps |
src/math_layout.rs (new) |
MATH-table loading + box layout | Medium; depends on font crate |
Cargo.toml |
Add read-fonts or ttf-parser for MATH table |
Low; both pure Rust |
pmacs-gpu/Cargo.toml |
Add Latin Modern Math font bundling | Low; ~200 KB compressed |
pmacs-gpu/src/main.rs |
Detect math spans in visible text, insert MathBox draw calls | Medium; touches main render path |
semantic_render.rs |
No changes (v0) | None by design |
Open questions
Q#IM1 — Font size and DPI scaling
Math glyph positioning includes the font size as an explicit parameter.
Inline math should match the base text font size; display math may use
a slightly larger size. How does the math layout engine receive the
current font size? Via the existing FontFacts protocol message
(protocol v17), or queried from the State fields directly?
Proposed: query self.font_size directly in the GPU frontend, same
as code text does. No protocol change. (Protocol is v18 at the time of
this framing — SUPPORTED=[6..=18].)
Q#IM2 — Color inheritance
Math glyphs should render in the foreground face color of the surrounding text, including themed faces. At minimum: math inside a string literal should inherit the string face color; math in comments should inherit the comment color. This implies the detection layer must cross-reference style spans at the boundary byte ranges.
Proposed (v0): render all math in the default foreground face. Color-by-context is deferred to v1.
Caveat: in markdown buffers, math inside a fenced code block should
render differently from math in prose. Since v0 defaults to foreground
color everywhere, math inside a ```rust block will be
indistinguishable from inline prose math. This is acceptable for v0 but
should be the first upgrade in v1.
Q#IM3 — Copy-paste fidelity
If a user copies a visual region containing a rendered \int_a^b,
what goes on the clipboard?
Proposed: the raw LaTeX source from the rope. This matches the
"math is text" invariant. A future refinement could copy the Unicode
math representation (e.g., ∫ₐᵇ) when that is available, but that is
a specialization for the clipboard protocol, not the rendering path.
Q#IM4 — Error fallback
What does an unparseable expression (unbalanced braces, unknown command) look like?
Proposed: render the raw LaTeX source with a red wavy underline
(reuses the existing diagnostic squiggle shader from
docs/pmacs-gpu-wavy-squiggles-framing.md). The expression is still
editable and copy-pasteable; the squiggle signals the parse failure
without losing the text.
Q#IM5 — Delimiter pairing visibility
$ and $$ are invisible delimiters in the rendered view. How does
the user know where the math region begins and ends when the cursor
is inside it?
Proposed: when the cursor is inside a math span, draw a subtle
background highlight over the entire range (reuses the Decorations
mechanism — MathFocus decoration produced by the GPU frontend
locally). The $ characters themselves are never hidden from the
underlying text buffer; they are merely suppressed in the glyph
pipeline. When the cursor approaches the boundary, the raw $
reappears.
Q#IM6 — Cursor navigation inside a math expression
The cursor moves through the raw LaTeX source at the byte level, not
through the rendered glyphs. The user edits \frac{a}{b} and sees it
render; the cursor steps through \, f, r, a, c, {, a,
}, {, b, }. The visual cursor position is a best-effort
projection: position the cursor at the x-offset of the rendered glyph
that corresponds to the nearest byte offset in the source.
Proposed: no structural cursor changes for v0. Cursor rendering
uses the existing CursorByte mechanism, which operates on the raw
text. The visual cursor may appear at a fractional screen position
inside a rendered expression; this is equivalent to how it appears in
an LSP inlay hint (already deployed — inline adornments hide their
source text and the cursor skips over them). V1 could smooth this.
Categorical bets
-
No JavaScript runtime. KaTeX is the best existing math renderer, but embedding a JS runtime (deno_core, quickjs, or a full V8) to run it would be the most expensive dependency pmacs has ever taken. Native Rust parsing + layout is ~2,000 lines vs ~30 MB of JS + WASM + runtime. The math layout algorithm is Knuth's 1978 design, well-documented and bounded in scope. Write it.
-
No SVG or pre-rasterization. Every expression renders fresh from the MATH table on every frame. This avoids a cache-invalidation explosion (what happens to a cached SVG when the user changes the font size? the theme? the DPI?) and keeps the rendering path uniform with code text (same glyphon pipeline, same GPU atlas).
-
The TUI is out for v0. Terminal cells cannot position glyphs at sub-cell resolution, cannot scale glyphs per-expression (except at great pain via Sixel / Kitty protocol), and cannot vary font size within a line. The TUI will render
$...$as-is with a distinct face (e.g., italic + a highlight color). This is honest: the feature is GPU-only, like the squiggle shader.Note: the TUI still needs tier-1 detection (to know which
$...$spans are math, so it can apply the distinct face). This makes detection a shared service — either the GPU frontend computes spans and the TUI queries them (impossible without a protocol message), or both frontends run their own detection. The v0 approach of frontend-local detection means the TUI must duplicate the delimiter scan. For v0 this is acceptable (the scan is cheap), but v1 should centralize detection in the instance and emitMathSpanson the wire. -
v0 is read-and-edit, not write-assist. No
\begin{align}completion, no auto-closing}, no preview of incomplete expressions. Those are UI conveniences layered on top once the pipeline exists.
Non-goals (explicitly excluded from this framing)
- Full LaTeX document rendering (titles, sections, bibliographies,
cross-references). That is a separate mode (
latex-mode.lua), not a frontend rendering concern. - MathML input. Detection parses
$...$/\(...\)syntax only. MathML→rendered would be a separate pipeline. - Real-time preview of incomplete expressions. A partially typed
\frac{will fail to parse and render as red-squiggled source; it does not show a partial fraction line. - Equation numbering and
\ref/\label. That is LaTeX-mode substrate, not math rendering.
Acceptance
All acceptance is GPU-side (the TUI renders raw $...$ and is not
separately tested for math). Tests run against pmacs-gpu with a
Vulkan device (PMACS_REQUIRE_GPU=1). Scratch buffers should clear
pmacs.lsp.config to avoid spurious server starts unless the test
exercises the LSP path.
-
Inline math renders: a buffer containing
The sum $\sum_{i=1}^n i$opens; the$...$range is suppressed in the glyph stream; a glyph run for∑, baseline-shiftedi=1andn, and the summation sign appears at the correct position between "The sum " and the following text. Visual assertion via GPU snapshot or caret/x-extent comparison. -
Display math centers: a buffer containing
$$\int_a^b f(x)\,dx$$opens; the math block occupies its own visual lines, centered horizontally, with vertical spacing above and below. -
Cache hit on re-scroll: scroll an expression off-screen and back; the second render takes < 1 µs (cache fetch). Measured via the GPU frontend's per-frame timing (already emitted to stderr).
-
Cache eviction on edit: edit inside a
$...$range; the expression re-parses (cache miss). Edit outside any math range; all cache entries survive. -
Error fallback renders as red-squiggled source: a buffer containing
$\frac{a$(unbalanced brace) opens; the raw source text\frac{a$is visible with a red wavy underline (reuses theSquiggleRendererpipeline). -
No false positive on bare
$: a buffer containingPrice: $5.00(single$with no pair) opens; no math spans are detected. The text renders as ordinary code text. -
Display math stores raw source for copy: select a visual region containing a rendered
$$\sum_{i=1}^\infty a_i$$; the clipboard receives the raw LaTeX source$$\sum_{i=1}^\infty a_i$$, not rendered glyphs. -
Cursor moves through raw source: cursor-right through
$\alpha$steps$,\,a,l,p,h,a,$(eight cursor positions) even though the visual rendering shows a singleαglyph. -
Stretchy fences assemble from glyph parts: a buffer containing
$\left(\frac{a}{b}\right)$renders the parentheses at least as tall as the fraction. Visual assertion (the parens enclose the fraction without clipping). -
Math survives buffer switch: open A (with math) and B (plain text), switch back; A's math cache is intact and no re-parse occurs.
Adjacent prior art in pmacs
-
Multi-language injections (
docs/multi-language-injections-framing.md): TheParseTreeBundle+Layermachinery already supports child parse trees at injected ranges. Math detection as a tree-sitter injection is a direct extension of this design, not a new mechanism. -
InlineAdornments (M11.3,
docs/semantic-frontend-protocol.md): LSP inlay hints demonstrate that the frontend can interleave virtual text at a byte offset without occupying document bytes. Math expressions need something strictly stronger — suppressing the raw$...$glyphs and replacing them with positioned math glyphs at potentially different widths. This suppression+replacement mechanism does not exist today (InlineAdornments are additive only —ChunkSource::SourceandChunkSource::Adornmentare interleaved, not exclusive). The math pipeline would build it, roughly following the same anchor/offset pattern but with aByteRangeto suppress and aMathBoxto render in its place. -
GPU squiggle shader (
docs/pmacs-gpu-wavy-squiggles-framing.md): Custom WGSL shader for diagnostic underlines demonstrates that pmacs-gpu already extends its rendering pipeline beyond basic glyph drawing. A stretchy-delimiter assembly pass or math-background pass follows the same architectural pattern (vertex buffer → dedicated shader → draw call in z-order). -
Theme-faces (Themes Arc 4,
docs/theme-faces-framing.md): The face resolution chain (face-attribute → color → fallback) would extend naturally to amath-faceormath-display-facefor themed math colors.