Review found the first version generating every envelope with an empty causal context. Production gives only the root that shape: each later edit carries active_prior_context(), the head's context extended by the head, so it covers the whole active prefix. Reduction orders on those edges, so a context-free log exercises a different algorithm -- and a cheaper one. Counters were wrong the same way: EditorSession mints at authored.len(), so the root is counter zero, which is also what extend_context recognises as the start of a contiguous run. Remeasured on a session-shaped log, reduce is roughly three times its former self at depth ten thousand -- 54 ms, not 17 -- and the wall moves from about ten thousand edits to between three and five thousand. That is the number T4b is sequenced against, so the first table would have mis-sequenced it. Depths three and five thousand now bracket the crossing; sampling only decades hid it. Two findings survive the correction and one is weakened. Reduce is still the only depth-scaling stage, and is superlinear at about n^1.4 -- which does not contradict the reduction bench's subquadratic result at fifty thousand envelopes, because that log is generated across three replicas with a different causal shape, and two logs of equal length are not equal work. Engrave is still flat, at 260 to 327 microseconds, and is still the larger half of the core's portion at shallow depth, so criterion 2's "uninformative while reduction dominates" holds only past roughly depth five hundred. But "render dominates at realistic depths" is now bounded: paint leads by four and a half times at depth one hundred, is level by one thousand, and is left behind after. T4 before T4b still stands -- the canvas removes what dominates a session's first thousand-odd edits -- but the two are no longer comfortably separated. The gate now includes envelope construction, which the requirement names first and the first version silently dropped. It is forty nanoseconds and never moves a verdict; a gate that omits a named component is a proxy for the requirement rather than the requirement. The 98% claim is replaced by both figures with their denominators named: what a direct-IR canvas avoids is 83% of the full measured per-edit pipeline, and 99.8% of the render path alone. The unqualified number was supported by neither. One row changed marking for a reason worth recording. Depth four thousand passes clean at 12.99 ms, but that is 78% of budget, and a load-contaminated run measured it at 22.77 ms -- above the five thousand row, which is impossible clean. A Pass row that fails whenever the machine is busy teaches people to ignore the gate, so the last gated Pass is three thousand and four thousand's clean number is kept as data in the table instead. Also recorded: depth is per session, not per document. EditorSession::open starts with an empty applied log, so reopening resets it and the reduced score becomes the new pristine base. That is what keeps a four-figure wall from being catastrophic -- though note entry mints one operation per note, so it is reachable in a sitting. Verified in an isolated worktree at HEAD rather than in the working tree, which currently carries the genesis tranche's in-flight G2 work: fmt clean, clippy 0 with and without golden-gate, workspace tests green, gate OK across all five rows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RSX4zSLgKvtiXaPjnMqLGz |
||
|---|---|---|
| .. | ||
| benches | ||
| examples | ||
| src | ||
| tests | ||
| Cargo.toml | ||
| DECISIONS.md | ||
| README.md | ||
README.md
epiphany-testkit
Agent F's crate per spec/QUICKSTART.md: the
cross-cutting conformance testkit. It is the architecture's tripwire — the suite
that proves the other crates work end to end and that runs in CI (see
.github/workflows/ci.yml).
It provides:
- Deterministic property-test generators for the public types of A
(
epiphany-determinism), B (epiphany-core), C (epiphany-ops), D (epiphany-bundle), and E (epiphany-layout-ir). Agent B's score-graph generators/shrinkers are re-exported asgenerators::graph. - The canonical round-trip harness (
roundtrip) — v0 acceptance criterion 4 (typed values + bundle container; the bookkeepingMaterializedStateround-trip is retained asassert_reduction_serialization_stable). - The CRDT convergence harness (
convergence) — criteria 1 and 5. Criterion 1 proper is real-Score convergence throughreduce_onto(run_graph_convergence); the byte-canonical bookkeeping-projection convergence (assert_convergence) backs criterion 5. - The equivocation harness (
equivocation) — criterion 3. - The crash-recovery harness (
bundle_harness) — Agent D's gate, criterion 2. - The manifest-selection harness (
bundle_harness). - The layout round-trip harness (
layout_stub) — criterion 6. - The audit regression guards (
negative) — one guard per defect the Agent C framework audit surfaced (the M1 fixes), so a regression trips this suite directly.
All harnesses are real
The QUICKSTART charters Agent F to "build against A and stubs for the others."
All five implementation crates — A, B, C, D, and now E (epiphany-layout-ir) —
have shipped, so every harness drives the real crate.
| Harness | Backend | Status |
|---|---|---|
roundtrip (criterion 4) |
A + B + C + D, real | real |
bundle_harness (criterion 2, manifest selection) |
D, real | real |
convergence (criteria 1, 5) |
C (epiphany-ops), real |
real |
equivocation (criterion 3) |
C (epiphany-ops), real |
real |
layout_stub (criterion 6) |
E (epiphany-layout-ir), real |
real |
For criteria 1, 3, and 5 the testkit drives the real
epiphany_ops::OperationSet / canonical_reduction_order / reduce and also
re-exports Agent C's own authoritative gates
(convergence::ops_reduction_determinism_fuzz,
equivocation::ops_equivocation_fuzz). The layout_stub module — once a
faithful in-tree stub of Chapters 7 & 9 — now re-exports the real
epiphany-layout-ir IR types and stub solver behind the same round_trip
signature; the provenance-preservation contract is implemented and tested inside
that crate. (The "stub" in the module name now refers to the spec-sanctioned
stub constraint solver, not to a stubbed crate.)
Criterion 1: real-Score vs. reducer-bookkeeping convergence
Criterion 1 proper (convergence::run_graph_convergence, the acceptance
criterion_1_convergence test) is real-Score convergence: a real ~50-bar,
two-voice base epiphany_core::Score is edited by two replicas through
OperationSet::reduce_onto, and the entire materialized graph — arena, voices,
tombstones, cross-cutting, and the bookkeeping state — must be identical
under every delivery order, pass check_invariants, and genuinely grow both
edited voices (non-vacuity). The session targets the base's actual voice ids
(generators::graph_edit_session), so it exercises the integration point, not a
synthetic id space.
The earlier, narrower gate is retained and honestly renamed
(reducer_bookkeeping_convergence): it converges the byte-canonical
bookkeeping projection (OperationSet::reduce →
MaterializedState::canonical_bytes) — the Chapter 6 §6.3 ledger (effects,
conflicts, anomalies, tombstones, spellings, pending), not the full musical
graph. It still backs criterion 5 and proves causal-first ordering
(convergence::assert_causal_order_respected,
run_authoritative_reduction_gate). The bookkeeping two-staff scenario remains
instantiated — a real ~50-bar (TWO_STAFF_BARS) session whose staves are
asserted populated by generators::assert_two_staff_populated, not just modeled.
Criterion 4: what is and isn't tested
Criterion 4 has three tiers — two asserted now, one pending item 5:
-
Real decode round-trips (these catch decoder / canonicalization defects): the generic
CanonicalEncode/CanonicalDecodeproperty swept across every typed identifier, bothRationalTimearms, and everyTypedObjectIddiscriminant; the bundleManifest(encode → decode → encodefixpoint, with a rich generator exercising snapshots, blobs, extensions, varied profiles, retention, and the optional roots); theFixedHeader; and theSuperblockslot encoding. Crucially, the decoders are shown to validate: corrupting a manifest or header makesdecodereject it (assert_manifest_decode_rejects_corruption,assert_header_decode_rejects_corruption). -
A reducer-bookkeeping serialization tier (
reducer_bookkeeping_serialization, viaassert_reduction_serialization_stable): a realOperationSetis reduced to itsMaterializedState::canonical_bytes()— the canonical bookkeeping state, not the whole musicalScore— which is stored as aSnapshotchunk referenced by the manifest'scanonical_base, survives the bundle's content-addressed store (hash-verified on reopen), decodes throughMaterializedState::decode_canonical, compares structurally with the original reduction, and re-serializes byte-identically. The decoder validates nested tags, lengths, primitive values, canonical form, and trailing bytes. Musical sensitivity is proven two ways:assert_content_mutation_changes_serialization(a cloned operation set with identical ids/stamps/causal contexts but one changed payload reduces to different bytes — the rebuttal to an id-only serializer) andassert_distinct_scores_serialize_differently. The materialized realScoreitself is shown reproducible today (full_score_materialization_is_reproducible, structural equality across delivery orders) — the determinism precondition a byte codec depends on. -
The full-
Scorebyte round-trip (criterion_4_full_score_byte_roundtrip, viaassert_score_serialization_stable): item 5's whole-score codec (epiphany_core::Score::canonical_bytes/decode_canonical) has landed, so a real ~50-barScore— materialized through Agent C'sreduce_onto— nowencode → decode → re-encodes byte-identically through a real bundle snapshot (hash-verified on reopen), with the decodedScorestructurally equal to the original. This is the whole musical graph (arena, voices, regions, cross-cutting, tombstones), not the bookkeeping projection.
Decisions (per QUICKSTART "Make each one once and document it")
- No platform entropy in the harness. Appendix D §"Randomness" forbids
platform entropy in canonical state; the testkit holds itself to the stronger
rule that no platform entropy enters the harness at all. Everything draws
from
rng::Rng(a wrapper over Agent A's vendored SplitMix64, with unbiased bounded draws and an overflow-safe full-rangerange), so every failure reproduces from its seed. - Drive the real crate once it ships; stub only what hasn't landed. Earlier
in development
epiphany-ops(C) andepiphany-layout-ir(E) were in-flight and their harnesses ran against faithful in-tree stubs; now that both have shipped, every harness drives the real crate and re-exports its gates.
Flagged for a future spec pass (Pass 11 candidates)
Per the QUICKSTART, implementation-discovered gaps are batched, not improvised:
- Whole-graph (
epiphany_core::Score) wire format — landed (item 5). A direct canonical byte codec for the coreScorenow exists (epiphany_core::Score::canonical_bytes/decode_canonical), andcriterion_4_full_score_byte_roundtripexercises it on a realreduce_ontomaterialization through a bundle snapshot. The prototype byte form predates the Binary Format companion specification and is to be reconciled with it (seeepiphany-core/DECISIONS.md, P11-4). - Layout harness re-pointed.
epiphany-layout-irhas landed, solayout_stubnow drives the real IR types behind the sameround_tripsignature (done). IR coordinates are f32 staff spaces, quantized only when serializing canonicalResolvedLayoutIR(Appendix D); see that crate'sDECISIONS.mdfor the remaining layout-specific Pass 11 candidates (theOperationKindTagvariant set and the layout-object id derivation).
Performance benches (Chapter 10 budgets, worklist F1)
benches/ holds the criterion benches for the spec's measurable Chapter 10
budgets (see DECISIONS.md F0 for why they live in this crate, F1 for every
call made). Criterion measures; the budget gate (src/budget.rs) asserts:
each bench's main() ends by re-timing every budget row and exiting nonzero if
a Pass-marked row misses its threshold. Known-pending rows are marked
Xfail(reason) in the bench source next to the numeric budget — a miss is
reported and tolerated, and a pass prints a loud promotion notice so stale
markings cannot linger. This is the "F surfaces, K fixes" handshake, and its
inaugural round has completed: the bench documented the reducer's O(n²)
canonical_reduction_order failure at scale, and Agent K's subquadratic
rewrite (see epiphany-ops/DECISIONS.md) flipped the xfail row to Pass.
| row | budget (spec Chapter 10) | expectation |
|---|---|---|
reduction/1000 |
> 10,000 envelopes/s, cold | Pass (~674K env/s measured) |
reduction/10000 |
> 10,000 envelopes/s, cold | Pass (~257K env/s measured) |
reduction/50000 |
> 10,000 envelopes/s, cold | Pass (~87K env/s measured; promoted from Xfail by Agent K's reducer fix — was ~1.7K env/s) |
bundle/typical_edit_commit |
≤ 50 ms (append + manifest + superblock flip, fsync'd) | Pass (~15 ms) |
bundle/open_bootstrap_read |
≤ 200 ms (manifest + bootstrap chunks) | Pass (moderate-corpus stand-in) |
# Full run (includes the 50K cold-reduction point, ~0.6 s per iteration):
cargo bench -p epiphany-testkit
# The reduced CI shape: smaller sampling, 50K point skipped (PR CI runs this):
EPIPHANY_BENCH_QUICK=1 cargo bench -p epiphany-testkit
The gate is a calibrated median over a few iterations, deliberately not the
spec's p99-over-1000-iterations conformance methodology (that is the reference
suite's job; the deviation is documented in src/budget.rs).
Running
# Unit + acceptance tests (a meaningful slice, under the cargo test timeout):
cargo test -p epiphany-testkit
# The full conformance suite at scale, outside the test timeout (includes Agent
# A's 1,000,000-iteration determinism gate and Agent C's reduction/equivocation
# gates):
cargo run --release -p epiphany-testkit --example conformance_suite # scale 1
cargo run --release -p epiphany-testkit --example conformance_suite 10 # soak
cargo run --release -p epiphany-testkit --example conformance_suite 0 # smoke
The six v0 acceptance criteria are asserted in tests/acceptance.rs, one test
per architecture layer.