epiphany/crates/epiphany-testkit
Levi Neuwirth 47fb4266c4 Ruling A criterion 2 stops being an assumption
The staged interactive-edit latency bench: reduce / engrave / scene-build /
paint measured separately, gating the core's portion against
req:perf:single-system-edit-latency's 16.7 ms frame. Criterion 2 asserts that a
toolkit verdict is uninformative while reduction dominates, and that sentence
had never been measured; the sequencing question it governs -- T4's spike now,
or T4b's incrementality first -- was resting on it.

The stage split is not invented here. It is the seam EditorSession::materialize
already walks, read off its private render_score and reproduced stage for
stage, so the bench measures the pipeline rather than a model of it. Only reduce
and engrave are gated: the requirement bounds "the core's portion" and says in
its own words that edit-to-pixel latency is a product-layer obligation, so
charging the SVG serializer and resvg against a core budget would be a category
error. They are measured and printed because the ruling asks for the stages
separately, and because today's is the path Ruling A demotes -- the number is
the baseline a canvas must beat, not a budget to defend.

Four findings, in the order they matter. Reduce is the only stage that scales
with log depth, near-linearly, and it breaks the frame at roughly ten thousand
edits -- 17.26 ms against 16.7, a three percent miss, so an order of magnitude
rather than a threshold. Engrave is flat and small at ~280 microseconds, and at
shallow depth it is the larger half of the core's portion, which qualifies
criterion 2 rather than confirming it: reduction does not dominate until about
depth five hundred. Paint is the largest single cost at every realistic depth --
2.12 ms at depth one hundred is four and a half times the entire core portion.
And scene-build is 3.5 microseconds of IR work plus about 130 of SVG
serialization, which the no-feature run separates: a canvas consuming the IR
directly skips some ninety-eight percent of today's per-edit cost, none of it in
the core.

The sequencing answer is therefore that T4 before T4b stands, for the opposite
reason to the one assumed. The dominant cost at the depths real sessions reach
is the render path Ruling A already demoted, not reduction. T4b's trigger is a
session ten thousand edits deep, and the bench now watches for it as the one
Xfail row.

Two things the bench had to survive being wrong about, both mine. The depth-1000
row was drafted Xfail on the assumption Fact 8 would already bite; it passes
with eightfold margin, the gate's XPASS notice said so, and the row is promoted
here rather than left stale -- which is the whole point of that mechanism. And
the first edit log alternated transposition direction per operation, which is
degenerate when the pitch-list length is even: every edit to a given pitch
pushed the same way, drifting it twenty-five semitones by depth 1000 and would
have been two hundred and fifty by depth 10000. That inflated engrave by a
factor of two and paint by nearly three -- a score-content change wearing a
log-depth costume. Alternating per pass instead bounds drift to one semitone.
The residual content effect is documented rather than hidden: paint is
non-monotonic in depth because pass-count parity decides how many accidentals
the score carries, and reading its dip at depth 10000 as a scaling win would be
a mistake.

Stated limitation: the testkit's largest fixture is three staves by ten
measures, so the engrave and scene-build columns are lower bounds and this
cannot prove the budget holds on the hundred-page orchestral score the
requirement contemplates. It shows where the time goes at the scale we can
build, and a row that misses at this size misses by more at a real one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz
2026-07-28 15:04:36 -04:00
..
benches Ruling A criterion 2 stops being an assumption 2026-07-28 15:04:36 -04:00
examples Editor T2 W3: the goldens become conformance gate [9/9] 2026-07-23 17:20:14 -04:00
src Editor T4-pre W1: the resolved layout stops discarding its own partition 2026-07-24 14:47:31 -04:00
tests Genesis G1: CreateInstrument, and the from-empty spine reaches a note 2026-07-24 21:02:12 -04:00
Cargo.toml Ruling A criterion 2 stops being an assumption 2026-07-28 15:04:36 -04:00
DECISIONS.md Editor T2 W3: the goldens become conformance gate [9/9] 2026-07-23 17:20:14 -04:00
README.md Push 4: Binary Format companion, F1 benches, subquadratic reduction order 2026-07-02 19:02:07 -04:00

README.md

epiphany-testkit

Agent F's crate per spec/QUICKSTART.md: the cross-cutting conformance testkit. It is the architecture's tripwire — the suite that proves the other crates work end to end and that runs in CI (see .github/workflows/ci.yml).

It provides:

  • Deterministic property-test generators for the public types of A (epiphany-determinism), B (epiphany-core), C (epiphany-ops), D (epiphany-bundle), and E (epiphany-layout-ir). Agent B's score-graph generators/shrinkers are re-exported as generators::graph.
  • The canonical round-trip harness (roundtrip) — v0 acceptance criterion 4 (typed values + bundle container; the bookkeeping MaterializedState round-trip is retained as assert_reduction_serialization_stable).
  • The CRDT convergence harness (convergence) — criteria 1 and 5. Criterion 1 proper is real-Score convergence through reduce_onto (run_graph_convergence); the byte-canonical bookkeeping-projection convergence (assert_convergence) backs criterion 5.
  • The equivocation harness (equivocation) — criterion 3.
  • The crash-recovery harness (bundle_harness) — Agent D's gate, criterion 2.
  • The manifest-selection harness (bundle_harness).
  • The layout round-trip harness (layout_stub) — criterion 6.
  • The audit regression guards (negative) — one guard per defect the Agent C framework audit surfaced (the M1 fixes), so a regression trips this suite directly.

All harnesses are real

The QUICKSTART charters Agent F to "build against A and stubs for the others." All five implementation crates — A, B, C, D, and now E (epiphany-layout-ir) — have shipped, so every harness drives the real crate.

Harness Backend Status
roundtrip (criterion 4) A + B + C + D, real real
bundle_harness (criterion 2, manifest selection) D, real real
convergence (criteria 1, 5) C (epiphany-ops), real real
equivocation (criterion 3) C (epiphany-ops), real real
layout_stub (criterion 6) E (epiphany-layout-ir), real real

For criteria 1, 3, and 5 the testkit drives the real epiphany_ops::OperationSet / canonical_reduction_order / reduce and also re-exports Agent C's own authoritative gates (convergence::ops_reduction_determinism_fuzz, equivocation::ops_equivocation_fuzz). The layout_stub module — once a faithful in-tree stub of Chapters 7 & 9 — now re-exports the real epiphany-layout-ir IR types and stub solver behind the same round_trip signature; the provenance-preservation contract is implemented and tested inside that crate. (The "stub" in the module name now refers to the spec-sanctioned stub constraint solver, not to a stubbed crate.)

Criterion 1: real-Score vs. reducer-bookkeeping convergence

Criterion 1 proper (convergence::run_graph_convergence, the acceptance criterion_1_convergence test) is real-Score convergence: a real ~50-bar, two-voice base epiphany_core::Score is edited by two replicas through OperationSet::reduce_onto, and the entire materialized graph — arena, voices, tombstones, cross-cutting, and the bookkeeping state — must be identical under every delivery order, pass check_invariants, and genuinely grow both edited voices (non-vacuity). The session targets the base's actual voice ids (generators::graph_edit_session), so it exercises the integration point, not a synthetic id space.

The earlier, narrower gate is retained and honestly renamed (reducer_bookkeeping_convergence): it converges the byte-canonical bookkeeping projection (OperationSet::reduceMaterializedState::canonical_bytes) — the Chapter 6 §6.3 ledger (effects, conflicts, anomalies, tombstones, spellings, pending), not the full musical graph. It still backs criterion 5 and proves causal-first ordering (convergence::assert_causal_order_respected, run_authoritative_reduction_gate). The bookkeeping two-staff scenario remains instantiated — a real ~50-bar (TWO_STAFF_BARS) session whose staves are asserted populated by generators::assert_two_staff_populated, not just modeled.

Criterion 4: what is and isn't tested

Criterion 4 has three tiers — two asserted now, one pending item 5:

  • Real decode round-trips (these catch decoder / canonicalization defects): the generic CanonicalEncode/CanonicalDecode property swept across every typed identifier, both RationalTime arms, and every TypedObjectId discriminant; the bundle Manifest (encode → decode → encode fixpoint, with a rich generator exercising snapshots, blobs, extensions, varied profiles, retention, and the optional roots); the FixedHeader; and the Superblock slot encoding. Crucially, the decoders are shown to validate: corrupting a manifest or header makes decode reject it (assert_manifest_decode_rejects_corruption, assert_header_decode_rejects_corruption).

  • A reducer-bookkeeping serialization tier (reducer_bookkeeping_serialization, via assert_reduction_serialization_stable): a real OperationSet is reduced to its MaterializedState::canonical_bytes() — the canonical bookkeeping state, not the whole musical Score — which is stored as a Snapshot chunk referenced by the manifest's canonical_base, survives the bundle's content-addressed store (hash-verified on reopen), decodes through MaterializedState::decode_canonical, compares structurally with the original reduction, and re-serializes byte-identically. The decoder validates nested tags, lengths, primitive values, canonical form, and trailing bytes. Musical sensitivity is proven two ways: assert_content_mutation_changes_serialization (a cloned operation set with identical ids/stamps/causal contexts but one changed payload reduces to different bytes — the rebuttal to an id-only serializer) and assert_distinct_scores_serialize_differently. The materialized real Score itself is shown reproducible today (full_score_materialization_is_reproducible, structural equality across delivery orders) — the determinism precondition a byte codec depends on.

  • The full-Score byte round-trip (criterion_4_full_score_byte_roundtrip, via assert_score_serialization_stable): item 5's whole-score codec (epiphany_core::Score::canonical_bytes / decode_canonical) has landed, so a real ~50-bar Score — materialized through Agent C's reduce_onto — now encode → decode → re-encodes byte-identically through a real bundle snapshot (hash-verified on reopen), with the decoded Score structurally equal to the original. This is the whole musical graph (arena, voices, regions, cross-cutting, tombstones), not the bookkeeping projection.

Decisions (per QUICKSTART "Make each one once and document it")

  1. No platform entropy in the harness. Appendix D §"Randomness" forbids platform entropy in canonical state; the testkit holds itself to the stronger rule that no platform entropy enters the harness at all. Everything draws from rng::Rng (a wrapper over Agent A's vendored SplitMix64, with unbiased bounded draws and an overflow-safe full-range range), so every failure reproduces from its seed.
  2. Drive the real crate once it ships; stub only what hasn't landed. Earlier in development epiphany-ops (C) and epiphany-layout-ir (E) were in-flight and their harnesses ran against faithful in-tree stubs; now that both have shipped, every harness drives the real crate and re-exports its gates.

Flagged for a future spec pass (Pass 11 candidates)

Per the QUICKSTART, implementation-discovered gaps are batched, not improvised:

  • Whole-graph (epiphany_core::Score) wire format — landed (item 5). A direct canonical byte codec for the core Score now exists (epiphany_core::Score::canonical_bytes / decode_canonical), and criterion_4_full_score_byte_roundtrip exercises it on a real reduce_onto materialization through a bundle snapshot. The prototype byte form predates the Binary Format companion specification and is to be reconciled with it (see epiphany-core/DECISIONS.md, P11-4).
  • Layout harness re-pointed. epiphany-layout-ir has landed, so layout_stub now drives the real IR types behind the same round_trip signature (done). IR coordinates are f32 staff spaces, quantized only when serializing canonical ResolvedLayoutIR (Appendix D); see that crate's DECISIONS.md for the remaining layout-specific Pass 11 candidates (the OperationKindTag variant set and the layout-object id derivation).

Performance benches (Chapter 10 budgets, worklist F1)

benches/ holds the criterion benches for the spec's measurable Chapter 10 budgets (see DECISIONS.md F0 for why they live in this crate, F1 for every call made). Criterion measures; the budget gate (src/budget.rs) asserts: each bench's main() ends by re-timing every budget row and exiting nonzero if a Pass-marked row misses its threshold. Known-pending rows are marked Xfail(reason) in the bench source next to the numeric budget — a miss is reported and tolerated, and a pass prints a loud promotion notice so stale markings cannot linger. This is the "F surfaces, K fixes" handshake, and its inaugural round has completed: the bench documented the reducer's O(n²) canonical_reduction_order failure at scale, and Agent K's subquadratic rewrite (see epiphany-ops/DECISIONS.md) flipped the xfail row to Pass.

row budget (spec Chapter 10) expectation
reduction/1000 > 10,000 envelopes/s, cold Pass (~674K env/s measured)
reduction/10000 > 10,000 envelopes/s, cold Pass (~257K env/s measured)
reduction/50000 > 10,000 envelopes/s, cold Pass (~87K env/s measured; promoted from Xfail by Agent K's reducer fix — was ~1.7K env/s)
bundle/typical_edit_commit ≤ 50 ms (append + manifest + superblock flip, fsync'd) Pass (~15 ms)
bundle/open_bootstrap_read ≤ 200 ms (manifest + bootstrap chunks) Pass (moderate-corpus stand-in)
# Full run (includes the 50K cold-reduction point, ~0.6 s per iteration):
cargo bench -p epiphany-testkit

# The reduced CI shape: smaller sampling, 50K point skipped (PR CI runs this):
EPIPHANY_BENCH_QUICK=1 cargo bench -p epiphany-testkit

The gate is a calibrated median over a few iterations, deliberately not the spec's p99-over-1000-iterations conformance methodology (that is the reference suite's job; the deviation is documented in src/budget.rs).

Running

# Unit + acceptance tests (a meaningful slice, under the cargo test timeout):
cargo test -p epiphany-testkit

# The full conformance suite at scale, outside the test timeout (includes Agent
# A's 1,000,000-iteration determinism gate and Agent C's reduction/equivocation
# gates):
cargo run --release -p epiphany-testkit --example conformance_suite        # scale 1
cargo run --release -p epiphany-testkit --example conformance_suite 10     # soak
cargo run --release -p epiphany-testkit --example conformance_suite 0      # smoke

The six v0 acceptance criteria are asserted in tests/acceptance.rs, one test per architecture layer.