Scope 6.6 deliverables 1-3, partial. This is evidence, not a freeze: no
Wave B deliverable set is complete and the throughput figure below is not
a publishable bundle.
The revert-and-observe-red acceptance for cd37f8b's two generators was
never actually performed -- the tests asserted it in their doc comments,
which is a claim. It is performed now, in a throwaway copy, one mutation
at a time, and each reddens exactly one test for the right reason. With
segment frame validation reverted to pushing footer offsets straight into
the adopted set, the store still opens and recovery still completes: it
reports recovery_ok=true and adopts the corrupted frame as authority, so
the red is exit 0 where 65 is required rather than a store that failed to
open. With the shard-index check removed from journal binding, the moved
journal is adopted whole and root_uuid alone does not catch it -- the Wave
A blocker reproduced.
The eight Wave B failpoint rows now drive StoreEngine::submit in-process
with the full six-field expectation asserted against the frozen oracle,
field by field rather than by struct comparison. No pending-wave-b row
remains. The four fields Wave A could not reach come from four independent
observations: the store naming itself poisoned, the two-root status read,
a receipt obtainable at all, and a different transaction submitted to the
same shard before any reopen. The last two are not the same question -- a
shard can be unpoisoned and still refuse a later append because its writer
thread died, which is exactly what phase-aware panic ownership fixed and
what this now checks from outside. AfterRootCasBeforeWaiterWake is
observed as a genuinely hung submit whose receipt is still retrievable,
so a waiter that never wakes is proved to be a hung request rather than an
absent transaction. Every row also asserts its group's fence count against
the public durability snapshot.
Two flake campaigns were measured rather than rerun: 4 failures in 40, then
7 in 40, from two distinct causes. One is a finding -- publishing a group
adds an index delta layer and none are sealed, so submit refuses after
exactly max_index_runs publications for the life of an engine. Both fixed
structurally; 200/200 and 40/40 after.
store-bench emit-skeleton now defaults to the submit path, with the journal
seam retained under --path drive for comparison. The signer is real, the
ref CAS is evaluated by the sequencer, and objects_new is summed from the
store's own receipts rather than multiplied out of the transaction count.
Explicit blockers, retained rather than worked around:
- StoreEngine::open still refuses startup state 1, so the benchmark seeds
its root by a non-production path. Seeding a store off the production
path in order to measure the production path is the charter item 8
smell; the disclosure is recorded in the fixture, a const doc, and the
module docs, and a test asserts open still refuses so it cannot go
stale in the safe direction.
- P2 is blocked three ways -- checkpointing disabled, no steady state
under the index-run ceiling, and no warmup/repetition/trim protocol.
The rate emitted is a debug build on tmpfs, marked preliminary.
- The 100 SIGKILL cycles still drive the journal seam, so kill -9 never
lands inside a real publication.
- Four schema claims became earnable and are requested, not emitted;
bench/result-schema.json is lead-owned.
- Reopen after close needs a bounded, measured, reported wait, because
the root LOCK outlives StoreEngine::drop. Diagnosed since as fork/exec
inheritance of the lock file description; the fix belongs in the lock
primitive, and this wait is removed when it lands.
scripts/check-phase1.sh GATE_EXIT=0; verify-store-recovery.sh 100 cycles,
recovery_failures=0, acknowledged_loss=0, torn_transactions=0,
repeated_adoptions=0, bundle=schema-valid.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A narrow vertical slice of scope 6.4: StoreEngine::open through one
retained RecoverySession, every shard recovered into the initial
CommittedRoot, and one shard writer wired GroupBuilder -> append_group_and_fence
-> ShardSubtree -> CAS publication -> completion. This closes the Wave A
carry-forward by giving GroupBuilder a production caller; the group bounds
are asserted against the sequencer as fdatasync counts on the device, not
against the builder in isolation.
§6.3's two load-bearing orderings are enforced and asserted, not assumed.
Step 8 follows step 7 with an explicit release fence, so a reader observing
the status root without an entry cannot then read a committed root older
than the publication -- verified by mutation in both directions. Waiters
are woken only after publication, and failure at steps 9-10 leaves a
queryable receipt, because a fence that succeeded and a root that published
is committed.
Three P1 findings from the B2 review are fixed here.
The committed event took its actor from the destination signer rather than
the evidence, so a frame could be signed, fenced, and published and then
fail frozen mirror verification -- first observed by another instance, long
after the bytes were durable. The test could not see it because the signer
key and the evidence actor were the same constant; they are now
deliberately different.
The deadline was rechecked in prepare, before signing and before the group
idle wait, so a slow signer could cross the retry deadline and still
append. The rule is now evaluated again over the whole group immediately
before anything is marked Resolving. The structural half matters more than
the recheck: shard_sequence is no longer consumed at prepare time, so a
dropped member leaves no hole by construction. Repair is re-derivation from
a recorded pre-image through the one sequencing function -- re-sequenced,
re-chained, re-signed -- because signing covers a digest that chains
previous_event_digest, and patching a suffix produces a durable, correctly
fenced frame whose signature verifies against nothing. The test reads every
frame back and checks both the chain and the signature; receipts alone
would not catch a partial repair.
Panic recovery had one catch around the whole writer loop, so every panic
poisoned the shard and reported every waiter as poisoned. Waiters now carry
an explicit phase and an owned reservation. A pre-append panic is
definitively absent and leaves the writer alive, since a dead writer makes
later_append_allowed_before_recovery false whatever the error says; a
request that overflows a publishing group is rolled back before that group
publishes rather than reported as its member; and a panic after publication
still delivers every receipt.
Also: the duplicated object-type and ref-kind tables are deleted in favour
of the amended shared helpers; max_objects_per_transaction and
max_refs_per_transaction are enforced at submit, the entry point a consumer
calls; status occupancy records the newly published root rather than the
one it replaced; and durability_counters/operation_status_metrics return
snapshots only, never the live counters, since a durability claim whose
auditor can write to it is not evidence.
StoreEngine::open constructs the single ProjectionStaging under the held
session before recovery -- it is recovery's projection resolver -- and
holds it exactly as long as the lock.
The remaining deliverables refuse by name: startup states 1/3/4,
coalescing, terminal retention, index sealing, journal rotation,
RepoSnapshot, and checkpoint. This is not a completed B1 deliverable set
and not freeze evidence; the gate's pending-row check stays inert while
engine.rs still returns NotImplemented.
138 library tests, 27 staging tests, 1 intentionally ignored.
scripts/check-phase1.sh GATE_EXIT=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scope 6.5 deliverables 1-5: what a staging session is on disk, what it
costs, when it dies. The instance-layer half of plan §8 — identity,
policy, ProjectionCore, the v2 routes — is deliberately absent; B3 binds
and exposes the fields those checks key on and evaluates none of them.
Two findings from the B2 review shaped the result more than the original
deliverables did.
Restart was in-memory handle reuse. open() built a fresh registry and
never read the filesystem, so after a real reopen an existing session ID
was admitted as new -- and if materialization then found the old
directory, error cleanup could unlink a durable session. That is data loss
reachable from an ordinary restart plus one error. open() now scans,
validates, and reconstructs sessions, chunk indexes, and the whole
occupancy account before returning, so a reconstructed ID is occupied and
refused as a conflict long before the error path; and that path no longer
calls remove_dir_all. Either change alone closes the loss.
A session is final iff its directory holds a valid session record for its
own ID, installed by rename_noreplace over fenced, digest-checked bytes,
so the final name only ever appears atomically over complete data. A
directory without one is abandoned materialization -- a crash between
mkdir and that rename, unaccounted and unreferenceable. The two states
share no code path, and reclaimed abandonments count on their own counter
so they can never be read as aborts or expiries.
Global quotas were per-handle. Every open() built an independent registry
outside the root LOCK, so two handles admitted twice the global limit and
the atomic-insertion work bought nothing across them. Construction now
requires proof of the held root lock and refuses a second in-process
instance, making two accountants on one root inexpressible rather than
discouraged.
Also: bounds are enforced atomically with insertion under one mutex with
no read-then-decide path, refused as typed LimitExceeded or Overloaded and
never by eviction; the directory-sync test pinned two syncs when the first
session in a shard needs three, a counter assertion that encoded the bug;
cleanup now validates a whole directory before unlinking anything, rather
than discovering a surprise midway through destroying a live session; and
artifact I/O moved off the registry mutex onto maintenance workers, with
the calling thread asserted to hold no guard rather than documented not to.
Deliverables 6-8 are explicit NotImplemented naming themselves.
StagedSessionState omits Finalizing, so deliverable 6 will fail to compile
at exactly the expiry and abort sites that must learn about a pin.
Carry-forwards recorded in §6.5, not closed: no production path begins a
session, so the sealed-invisibility acceptance stays ignored with both
blockers named; and expire() has no scheduler, so session age is a bound
enforced when asked and never asked.
116 library tests, 27 staging tests, 1 intentionally ignored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The B2 review of the first B1/B3 slice found five defects whose fixes were
not available to the packages that had to make them: each needed a change
to a surface those packages do not own. Rather than let them restate a
frozen fact locally, the lead amends the surfaces and the packages consume
them. Contract review 2026-07-28-A records all five with the weaker
alternative that was rejected for each.
- CommitEvidenceSigner::sign_event's parameter becomes signing_digest.
The message is SignedCommittedTransactionV1::signing_digest, which
binds key epoch and durability result around the canonical event; the
bare event digest is chain identity, not signature material. A doc line
was not enough: misuse is undetectable until after durability, and is
first observed by a mirror on another instance.
- TransactionEvidenceV1::actor() is new, exhaustive over all six
variants. A destination event's actor must restate the evidence's
source instance, and deriving it a second way store-side is what
produced the defect it closes.
- format::object_type_code becomes pub(crate) and recovery.rs's private
twin is deleted, so one table exists where three did.
- RefRecord gains from_target/target() in terms of RefTarget, with the
code table stated once per direction and an unknown kind refused by
value as CheckpointError::RefKind. Defaulting an unknown kind would
launder it into the next checkpoint within one interval.
- max_projection_objects is capped at MAX_CANONICAL_ITEMS and its default
lowered to it; max_projection_chunks likewise, and max_projection_bytes
against the transitive per-chunk allowance. The old default described a
projection no manifest could encode. Supporting a hundred million
objects requires a versioned chunked or indexed manifest design, not a
larger hostile-decode ceiling.
Both governing documents are updated where they now misstate a frozen
fact, including §6.4's claim that signing covers the event digest. §6.5
records the two carry-forwards from B3's first slice — production staging
use is incomplete, and expire() has no scheduler — and records that the
requested StoreEngine staging accessor is a pending amendment which must
not be a bare Arc<ProjectionStaging>, since a clone could outlive
EngineShared, survive release of the root LOCK, and keep serving a root
this process no longer holds.
No package file is touched: engine.rs, transaction.rs, and staging.rs
consume these in their own commits.
levcs-protocol 6 consumer tests, levcs-store 109 library tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Wave A review recorded that the crash matrix structurally could not
express its own first blocker: no failpoint corrupts a frame inside a
sealed segment, and none moves a journal between shards of one root. A
green matrix on a defect it cannot represent is the same trap as a helper
nothing calls, so the finding was carried forward rather than closed.
Add the two physical crash-image generators as a `damage` subcommand on
the crash driver, and two tests that drive them through the same
production `reconcile` path the matrix and the recovery script use.
- sealed-frame-corruption flips one payload byte in the first frame of
a segment the drive actually sealed and installed. The footer stays
structurally valid, so only frame verification can reject it.
- cross-shard-journal-movement relocates shard 1's active journal under
shard 0 of the same root. Root UUID validation cannot see this; the
journal header's shard index must be bound to the directory being
opened.
Both generators mutate a production-written image and fence the mutation;
neither synthesizes a footer, frame, manifest, or checksum, so a passing
test cannot be an artifact of the harness agreeing with itself. Each
source image is required to be unambiguous — exactly one segment or
journal — so the result does not depend on directory iteration order.
The assertions pin the refusal to its own cause rather than to any
non-zero exit: the frame case must fail on the digest recomputation and
the movement case on the shard binding, and neither may publish a partial
adoption result.
crash_matrix: 26 passed. scripts/check-phase1.sh green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The freeze commit cannot contain its own hash, so it is recorded here.
Wave A is frozen at 5ee9c6b78b with the gate
evidence measured at that commit, in both the plan and the scope document.
Restates what the freeze does and does not cover: the two carry-forwards are
B-wave work and outside it, and doc/swarm-fabric-roadmap-exploration.md
(e6a058d) stays non-binding and outside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG
Deliberately a separate commit, outside the Wave A freeze (5ee9c6b). This
document decides nothing and is not part of any frozen contract: it is a place
to hold the thinking on whether LeVCS should become a change fabric for
multi-agent work, pending its own review.
Records where the proposal maps onto the existing plan, where it cuts against
it -- levcsd must not own ref transactions, the ephemeral stratum invariant,
delegation being identity-shaped rather than protocol-shaped, attestations not
fitting TransactionEvidenceV1, and the canonical benchmark workload being the
wrong shape for swarm traffic -- and where the proposal is simply wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG
Freezes the Wave A interfaces, frame format, durability ordering, and
crash/fault fixtures per doc/phase1-storage-spine-scope.md section 5. Wave B
may now start against these.
D0 lands crates/levcs-store: the sealed public API, the file-ownership split,
the single durability syscall funnel with its counters and fault hooks, the
17-entry failpoint registry in compiler-enforced correspondence with
oracle::AppendFailpoint, and the journal-level drive seam. Wave A lands the
frame codec and journal/segment lifecycle (A1), the recovery index and
checkpoints (A2), and the durable-ingest benchmark and crash harness (A3).
The adversarial review found five defects behind a green gate, three of them
blockers, all closed here. Two were the same shape: drive::reopen_through_
recovery had reimplemented a simplified recovery and called none of recovery.rs
-- so it adopted segment footer sequences without validating frame bytes,
journal_id, or root_uuid, and it double-adopted interrupted seals. The seam
between two packages was untested precisely because each package's own tests
passed. reopen_through_recovery now delegates rather than decides, and
DriveRecovery carries the recovery report verbatim so tests assert the
disposition and not merely its effect: a double adoption and a correct replay
produce the same adopted set, which is how the defect stayed invisible.
Also closed: store-bench now schema-validates its own emitted artifact with a
five-mutation negative control instead of matching JSON substrings; the ACK
reconciler distinguishes duplicates and regressions from forward gaps;
checkpoint writes and a journal truncation are routed through the durability
funnel, whose guard now covers writes and truncations rather than only sync,
rename, and unlink.
bench/result-schema.json is amended (contract review 2026-07-24-B, second and
third amendments): per-gate latency ceilings conditional on outcome so a failed
run is representable, and the verification claims split per gate so a storage
run cannot certify an object graph it never touches. Not-applicable claims are
forbidden rather than falsified; applicable-but-not-performed report false.
Every relaxation is re-pinned in the else branch and asserted member by member,
after an edit in this series silently un-pinned all eleven validation flags and
was caught only by revalidating against constructed bundles.
Arming the fault registry now requires a FaultSerial token, so the invariant is
a compile error rather than a comment. The file where this was diagnosed
carried a header saying it was deliberately the only test in it, and a second
test had been added under that comment anyway -- an 8-in-40 failure rate that
read as flakiness.
Evidence at this commit: check-phase1.sh GATE_EXIT=0 across all four feature
configurations, 124 test binaries, zero failures; verify-store-recovery.sh
--cycles 100 with recovery_failures=0, acknowledged_loss=0,
torn_transactions=0, repeated_adoptions=0, bundle=schema-valid; recovery_eio
40/40 at four test threads; golden corpus byte-stable; fmt clean.
Carry-forwards, explicitly not Wave A blockers and recorded in scope section 5:
extend the crash matrix to generate sealed-frame corruption and cross-shard
journal movement, since it structurally cannot express the class the first
blocker belonged to; and wire GroupBuilder through B1's production path, since
deliverable 4-A1.2 is presently asserted only over a type nothing calls.
Charter item 9 -- ask every package what of its work is correct but uncalled --
is accepted for every subsequent wave. It, not the review, found the class.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG