Commit Graph

9 Commits

Author SHA1 Message Date
Levi Neuwirth b724e94faf Contract revision 6: round 1 was testing the wrong property on the wrong glyphs
Building round 1's oracle exposed two defects in the round as written, both
mine, and both invisible until something tried to satisfy it.

The glyph set was chosen by subpath count. I wrote that 19 of the 37 bundled
outlines have more than one subpath and named gClef, fClef, timeSig8 and
accidentalFlat as the useful ones -- conflating multi-subpath with has-a-hole.
fClef is the counter-example: its three subpaths are a bowl and two SOLID,
disjoint dots, nested in nothing. Measured across all 37 by point-in-path,
exactly twelve carry a bounded hole; fClef, cClef, barlineFinal and every
repeat glyph carry none.

Round 1 now runs five glyphs in two classes testing two different properties.
Hole checks -- gClef, timeSig8, accidentalFlat and noteheadHalf, which is new
and earns its place by being frequently repeated and semantically consequential
(a filled counter renders half notes as quarter notes, a notation error rather
than an artifact). Disjoint-component check -- fClef alone, no background
requirement, instead requiring one ink point inside EACH of its three filled
subpaths, tagged with its subpath index. That second class catches a
tessellator that keeps only the largest contour, which would pass every hole
check ever written. Hard failure is now stated as either: a bounded hole
painted as ink, or a required filled subpath omitted.

The oracle's status model is now specified rather than inferred from an
absence. fClef passing with zero background points is a SATISFIED result under
its own requirement class; recording it only as background_satisfied = false
would make a correct outcome indistinguishable from a failed one.

The second defect was the criterion itself. Ruling A said epaint does not
implement even-odd/nonzero fill for paths with holes -- framing criterion 1
around the fill RULE. Bravura's contours are correctly oppositely wound
(signed ring areas gClef [8.702, -0.691, -1.803, -0.509]; fClef
[2.534, 0.153, 0.148], all positive, the same fact from the other side), so
even-odd and nonzero AGREE on every bundled hole. The rule is not load-bearing;
preserving every filled contour and every bounded counter is. The criterion is
amended to compound-path / inner-subpath fill correctness, with a more accurate
reason for excluding raw egui shapes than the one it replaces: PathShape is a
single point loop documenting "Fill is only supported for convex polygons", so
it cannot express compound-fill or subtractive-hole semantics -- a Shape::Vec
can group loops, but grouping paints them, it does not subtract a counter from
its enclosing contour.

Signed areas are recorded from the oracle's adaptive flattening rather than an
earlier coarse fixed-step measurement, with the note that magnitudes are
flattening-dependent and the SIGNS are the claim.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 21:24:47 -04:00
Levi Neuwirth 4a4988ce88 The latency wall is document-lifetime, not one sitting
Review found the mitigation attached to the last landing reading the wrong
code path. EditorSession::open does start with an empty applied log, but that
is the score-only probe constructor. The savable-document path is Ruling B's:
reopen is full replay, stored envelopes load as a committed partition, and
materialization reduces committed plus session operations together. Nothing
resets the depth this bench varies until the checkpoint and pruning machinery
assigned to T4b can write a new canonical_base.

So consequence (e) is withdrawn rather than corrected in place. The ~4,500-edit
wall is a budget on a document's whole accumulated history, and there is no
session reset to lean on -- neither as reassurance about the number nor as
support for the sequencing argument, which rests on paint dominance and does
not need it. T4b's trigger is correspondingly firmer than it read yesterday.

Also: two comments still described the gated core portion as reduce plus
engrave, from before envelope construction was added as a third stage. The
sum they document has included it since the last landing.

Verified in an isolated worktree at HEAD rather than in the working tree,
which still carries the genesis tranche's in-flight work: fmt clean, clippy 0
with and without golden-gate, gate OK with every verdict unchanged (the edits
are documentation only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 16:27:59 -04:00
Levi Neuwirth 986c9cc5a1 The edit-latency bench measures the log a session actually writes
Review found the first version generating every envelope with an empty causal
context. Production gives only the root that shape: each later edit carries
active_prior_context(), the head's context extended by the head, so it covers
the whole active prefix. Reduction orders on those edges, so a context-free log
exercises a different algorithm -- and a cheaper one. Counters were wrong the
same way: EditorSession mints at authored.len(), so the root is counter zero,
which is also what extend_context recognises as the start of a contiguous run.

Remeasured on a session-shaped log, reduce is roughly three times its former
self at depth ten thousand -- 54 ms, not 17 -- and the wall moves from about ten
thousand edits to between three and five thousand. That is the number T4b is
sequenced against, so the first table would have mis-sequenced it. Depths three
and five thousand now bracket the crossing; sampling only decades hid it.

Two findings survive the correction and one is weakened. Reduce is still the
only depth-scaling stage, and is superlinear at about n^1.4 -- which does not
contradict the reduction bench's subquadratic result at fifty thousand
envelopes, because that log is generated across three replicas with a different
causal shape, and two logs of equal length are not equal work. Engrave is still
flat, at 260 to 327 microseconds, and is still the larger half of the core's
portion at shallow depth, so criterion 2's "uninformative while reduction
dominates" holds only past roughly depth five hundred. But "render dominates at
realistic depths" is now bounded: paint leads by four and a half times at depth
one hundred, is level by one thousand, and is left behind after. T4 before T4b
still stands -- the canvas removes what dominates a session's first thousand-odd
edits -- but the two are no longer comfortably separated.

The gate now includes envelope construction, which the requirement names first
and the first version silently dropped. It is forty nanoseconds and never moves
a verdict; a gate that omits a named component is a proxy for the requirement
rather than the requirement.

The 98% claim is replaced by both figures with their denominators named: what a
direct-IR canvas avoids is 83% of the full measured per-edit pipeline, and 99.8%
of the render path alone. The unqualified number was supported by neither.

One row changed marking for a reason worth recording. Depth four thousand passes
clean at 12.99 ms, but that is 78% of budget, and a load-contaminated run
measured it at 22.77 ms -- above the five thousand row, which is impossible
clean. A Pass row that fails whenever the machine is busy teaches people to
ignore the gate, so the last gated Pass is three thousand and four thousand's
clean number is kept as data in the table instead.

Also recorded: depth is per session, not per document. EditorSession::open
starts with an empty applied log, so reopening resets it and the reduced score
becomes the new pristine base. That is what keeps a four-figure wall from being
catastrophic -- though note entry mints one operation per note, so it is
reachable in a sitting.

Verified in an isolated worktree at HEAD rather than in the working tree, which
currently carries the genesis tranche's in-flight G2 work: fmt clean, clippy 0
with and without golden-gate, workspace tests green, gate OK across all five
rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 16:03:27 -04:00
Levi Neuwirth 47fb4266c4 Ruling A criterion 2 stops being an assumption
The staged interactive-edit latency bench: reduce / engrave / scene-build /
paint measured separately, gating the core's portion against
req:perf:single-system-edit-latency's 16.7 ms frame. Criterion 2 asserts that a
toolkit verdict is uninformative while reduction dominates, and that sentence
had never been measured; the sequencing question it governs -- T4's spike now,
or T4b's incrementality first -- was resting on it.

The stage split is not invented here. It is the seam EditorSession::materialize
already walks, read off its private render_score and reproduced stage for
stage, so the bench measures the pipeline rather than a model of it. Only reduce
and engrave are gated: the requirement bounds "the core's portion" and says in
its own words that edit-to-pixel latency is a product-layer obligation, so
charging the SVG serializer and resvg against a core budget would be a category
error. They are measured and printed because the ruling asks for the stages
separately, and because today's is the path Ruling A demotes -- the number is
the baseline a canvas must beat, not a budget to defend.

Four findings, in the order they matter. Reduce is the only stage that scales
with log depth, near-linearly, and it breaks the frame at roughly ten thousand
edits -- 17.26 ms against 16.7, a three percent miss, so an order of magnitude
rather than a threshold. Engrave is flat and small at ~280 microseconds, and at
shallow depth it is the larger half of the core's portion, which qualifies
criterion 2 rather than confirming it: reduction does not dominate until about
depth five hundred. Paint is the largest single cost at every realistic depth --
2.12 ms at depth one hundred is four and a half times the entire core portion.
And scene-build is 3.5 microseconds of IR work plus about 130 of SVG
serialization, which the no-feature run separates: a canvas consuming the IR
directly skips some ninety-eight percent of today's per-edit cost, none of it in
the core.

The sequencing answer is therefore that T4 before T4b stands, for the opposite
reason to the one assumed. The dominant cost at the depths real sessions reach
is the render path Ruling A already demoted, not reduction. T4b's trigger is a
session ten thousand edits deep, and the bench now watches for it as the one
Xfail row.

Two things the bench had to survive being wrong about, both mine. The depth-1000
row was drafted Xfail on the assumption Fact 8 would already bite; it passes
with eightfold margin, the gate's XPASS notice said so, and the row is promoted
here rather than left stale -- which is the whole point of that mechanism. And
the first edit log alternated transposition direction per operation, which is
degenerate when the pitch-list length is even: every edit to a given pitch
pushed the same way, drifting it twenty-five semitones by depth 1000 and would
have been two hundred and fifty by depth 10000. That inflated engrave by a
factor of two and paint by nearly three -- a score-content change wearing a
log-depth costume. Alternating per pass instead bounds drift to one semitone.
The residual content effect is documented rather than hidden: paint is
non-monotonic in depth because pass-count parity decides how many accidentals
the score carries, and reading its dip at depth 10000 as a scaling win would be
a mistake.

Stated limitation: the testkit's largest fixture is three staves by ten
measures, so the engrave and scene-build columns are lower bounds and this
cannot prove the budget holds on the hundred-page orchestral score the
requirement contemplates. It shows where the time goes at the scale we can
build, and a row that misses at this size misses by more at a real one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 15:04:36 -04:00
Levi Neuwirth f639919ee1 Editor T4-pre W3: the text run carries both its string and its ink
The last T4 prerequisite, and the one that had to be a ruling rather than a
packet: a canvas, an exporter, a hit test, and an accessibility tree must agree
about text, and what they agree on is decided by where the shaper runs relative
to the canonical boundary.

The census reframes the tranche. Three of the five categories the plan names --
lyrics, chord symbols, rehearsal marks -- carry no text in the model at all:
LyricLine holds only event references, ChordSymbol and Marker only an anchor.
They are blocked on a core-track schema major, not on this decision. What the
primitive does gate is the text the model already has: score metadata,
instrument and staff names, and text-line spanners. That is a smaller v1 than
the plan implied and a real one, and bidi and fallback are exercised through
synthetic fixtures that need no model work.

The ruling is a fourth resolved primitive carrying the source string and the
canonical shaped result together. The alternative that discards the string
renders a title as anonymous outlines and is unreadable to a screen reader; the
alternative that discards the shaped result lets two consumers draw the same
bytes differently, which contradicts the definition of canonical_bytes as the
rendering fingerprint. Both halves stay, and the apparent trade between
deterministic geometry and accessibility turns out not to exist.

Two drafts were wrong in opposite directions and the errors are recorded rather
than quietly fixed, because each came from asserting a constraint instead of
reading the requirement that governs it. Draft 1 held that shaping before the
canonical boundary poisons cross-implementation byte equality -- but layout
determinism is byte-equal only within one implementation at a fixed version;
across implementations it is reference-suite thresholds, and the spec says so in
both the determinism table and req:solver:cross-implementation-conformance.
Draft 1 had imported the score layer's guarantee into the layout layer, where
the spec deliberately weakens it. Revision 2 then over-corrected, banning host
fonts outright on the grounds that an OS font update breaks fixed-version
stability -- but that requirement defines identical inputs to include font
metrics referenced by version and content hash, so an updated font is a changed
input. The rule that survives is narrower than either: no ambient or unresolved
lookup, and a host face may participate only once resolved to an exact
content-hashed asset every consumer can obtain.

The identity is specified rather than gestured at, because bytes that do not
determine ink are worse than bytes that admit they don't. A face is pinned by a
hash over the font file, not its metrics -- GlyphCatalogIdentity's metrics_hash
covers bounding boxes, advances and anchors, which pins spacing and not shape --
together with face index, variation coordinates and synthetic weight/slant.
Segments carry font-internal glyph ids, source ranges, direction, script,
language and em size; glyph offsets have alignment already applied, so a
consumer places by origin alone; positions quantize on the same 1/1024 grid as
every other primitive. The cluster map indexes UTF-8 byte offsets with caret
stops at grapheme boundaries carrying bidi affinity, and the Unicode
segmentation version is always part of the identity -- otherwise two
implementations could agree on every pixel and still differ inside the
fingerprint, where no visual test would ever see it.

One consequence lands on the exporter: SVG cannot honour "no consumer reshapes"
with <text>, which carries characters and lets the viewer's shaper choose the
glyphs, so a ligature or positional form silently draws something the layout did
not resolve. Conformant text export emits explicit glyphs as paths through the
same face, reusing the mode render-svg already has for music.

The reservation is re-ordered to follow shaping rather than precede it -- with a
canonical shaper in the pipeline, reserved_box becomes a solver policy over
measured bounds, not an estimate of them. Paint-time re-spacing stays forbidden.

Two findings for the core track, named so their absence is a decision. Score
text authored through operations is not NFC-validated: the envelope's NFC-
checked string reader covers only directly encoded strings such as transaction
labels, while SetMetadataOp, CreateStaffOp, CreateInstrumentOp and the cross-
cutting values embed the core codec's bytes, which preserve non-NFC strings by
design. And the .tex amendment adding the primitive changes the layout
fingerprint but needs no bundle or wire schema-major move, following strokes and
curves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 14:38:06 -04:00
Levi Neuwirth a35f534235 Editor T4-pre W2: the glyph-asset contract, and three parallelism claims get finer
The W2 contract charters the shared typed glyph-asset seam Ruling A names as a
T4 prerequisite. Scoping it turned up the same shape W1 had: the seam is
already designed and merely unpopulated. PathCommand, GlyphRenderData, and
GlyphCatalog::render_data all exist in layout-ir, and BravuraCatalog returns
None by deliberate documented honesty -- reporting Some would claim render data
that does not exist. So the packet fills a seam rather than building one.

Two findings reshape it from the sketch carried in the W1 contract. Bravura.otf
is not in the tree -- tools/ holds only the extractor script and OFL.txt, and
the generated header pins source hashes verified at extraction time -- so
"have the generator emit typed paths alongside the d strings" cannot be
executed here. The contract replaces it with a dependency-free in-crate parser
over exactly the grammar the generator emits, and proves equivalence by
round-trip: parse every bundled d, re-emit, compare byte-for-byte, with a
sanctioned coordinate-sequence fallback that must be reported if used. And the
metrics table is conformance identity -- metrics_hash hashes (name, metrics)
pairs with values participating, and GlyphCatalogIdentity is encoded into the
resolved layout's canonical bytes -- so it is out of bounds entirely.

The test worth watching is the cross-table one. glyph.rs claims the metrics and
the outlines agree because both came from the same Bravura release; that is
asserted in prose and tested nowhere. The contract requires comparing each
glyph's real outline extent against its declared bbox, reporting the worst-case
deviation, and treating a failure as a finding rather than a reason to widen
the tolerance -- it would mean engraving reserves the wrong space for that
glyph, which is the bug class that twice bit the vertical metric.

Three parallelism claims are corrected in the same pass, all mine and all too
coarse. The plan and the ruling both said T1b's lease/save/single-writer
machinery could be contracted in parallel with the genesis work. That was
written before the ladder existed, and it is now per-rung rather than
unconditional: G1 needs no accept-set raise and never enters epiphany-bundle,
so T1b's bundle work runs beside it, while G2 spends the raise in bundle.rs
where T1b's single-writer enforcement also lands, so those two must not fly
together. Ruling B's blocker note also still described the identity disposition
as blocking; it is ruled, and now points at where.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-24 17:50:02 -04:00
Levi Neuwirth 011c68a831 The operation set absorbs genesis
Ratified 2026-07-24, resolving Ruling B blocker (i). Pass-12 K8 is reversed:
every mutable field of Score becomes operation-authored, and a document is
Score::empty(identity) plus its envelope log. The instrument, staff, staff
instance, voice, event chain is authorable end to end.

The decision removes machinery rather than adding it. Every alternative kept
genesis outside the operation set and then had to pay for that: a new chunk
role, a manifest field, an immutability rule, and a merge or fail-closed rule
for a canonical payload with no CRDT semantics. Genesis state is edited —
instruments get added, page geometry changes, temperaments are chosen — and
each alternative made those edits single-writer, unmergeable, or impossible.
Concurrency is a first-order product commitment, so the exception was not
worth institutionalising in the format.

Scope is nine surfaces over two templates already proven in reduce.rs: three
LWW settings setters on the SetMetadata pattern (canvas.layout_defaults,
tuning_context, spelling_precedence) and six entity mint families on the
CreateStaff pattern (instruments, staff_groups, parts, analysis_layers, views,
and StaffInstance.measures), each with graph-aware referential preconditions.
Delete and modify coverage is left to the tranche contract rather than assumed,
since CreateStaff itself ships today with no DeleteStaff.

Measures are ruled authored rather than derived. TimeAnchor::Measure carries a
measure id that cross-cutting structures anchor to, so deriving measures from
the metric grid would make their identity a function of the meter and every
time-signature change would orphan the anchors pointing into them. The cost
accepted is that measure/meter consistency becomes an authoring obligation
backed by a graph invariant.

Three constraints are written in rather than left implicit. Pruning may not be
implemented until the canonical base carries graph values: a prune installs a
MaterializedState base whose effects are outcomes, not payloads, so nothing
rebuilds the score afterward — silent and total, and free to prohibit now
because no prune exists to break. The from-empty path must reduce through
new_onto with an empty Score rather than base-free, because the base-free mode
skips graph-aware preconditions by design and would silently lose referential
enforcement from the first operation. And the OperationEnvelopeBlock accept-set
raise 2 to 3 is spent once, so the new kinds land as one batch — this is a
different major from Push 4b's schema major 3, the Score and Snapshot role wire
that tranche 3b-i froze, and there is no free ride between them.

The analysis is corrected in place rather than rewritten, so the evidence the
ruling rests on stays readable. Two amendments: Measure is a ninth uncovered
surface the original table missed by scoring canvas.regions at container
granularity, and identity is promoted from a stated question to a blocking one
— IdentityContext is replica-scoped yet lives on Score and is encoded, so under
from-empty reduction two replicas with an identical log produce Scores
differing in an encoded field while the music is identical. That disposition
blocks specification of the tranche and is deliberately not ruled here.

Execution belongs to the Push-4b-class coordinated track; the editor track
consumes it. T1b's lease, save, and single-writer machinery does not depend on
the tranche landing and may be contracted in parallel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-24 15:11:36 -04:00
Levi Neuwirth 35bba767d2 Editor T2: contract for selection v2, the golden gate, and copy/paste
Four work packets: the selection set with an anchor (W1), GUI rubber-band
select (W2), promotion of the T1a goldens to conformance gate [9/9] behind
a golden-gate feature that keeps resvg out of the MSRV closure (W3), and
copy/paste over the newly granted Ruling E fragment projection (W4) —
values-only, paste-as-minting, fail-closed closure, untrusted-input caps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 17:07:18 -04:00
Levi Neuwirth 0e59f545ba Editor track: the plan, and the T1a golden-harness contract
PLAN_EDITOR_APP.md charters the editor product track: rulings A/C granted,
B blocked behind graph-state-persistence + versioned-decode, D conditional
on the document-bound session API; hardened by three source-level reviews
(14 + 11 + 9 findings, all dispositioned in its ledgers).
CONTRACT_EDITOR_T1A_GOLDENS.md dispatches the first tranche: pixel goldens
over the score raster, subagent work packets, coordinator review, user
deep-dive points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 14:57:43 -04:00