Review of 2026-07-30-A declined the silent state it disclosed. When a `.seg`
occupies the logical generation the active tail carries, the frames must be
renamed, and the two cases part on what that costs:
- orphan alone: nothing names the displaced identity, the seal falls back to
a free generation, and recovery succeeds — an orphan segment is a state in
which the active journal is still the authority, and refusing would turn a
recoverable root into an outage;
- orphan plus a run the manifest names: recovery would publish a manifest
naming a run whose every location resolves to nothing. It refuses with
Corruption instead, naming the run.
Refusing is an outage on a root whose data is all present. It is chosen because
the alternative opens and lies, and because the replay delta above the run hides
that from every lookup until the first consumer that reads runs directly — which
is the checkpointer this precedes.
`IndexRun::references_segment_generation` is the frozen seam, read-only, with
recovery as its one caller (contract review 2026-07-30-B). Exact rather than a
range test over section headers: entries pack a 16-bit delta from the section
base, so the header says only what a section could name, and a `true` it does
not owe refuses a recovery with nothing to lose.
`coverable_through` keeps the coverage 2026-07-30-A widened — the unsound state
is now refused where it arises rather than designed around at every seal — and
its comment, which still described the pre-split world, says so.
Evidence: the refusing test asserts the damage rather than an expectation, so
disabling the guard reports that the reopened root pins None at the generation
the run names. A first draft of it passed for the wrong reason, re-pushing the
genesis object id as a blob so the reopen failed on a duplicate-object Conflict
either way; the mutation exposed that.
Still open and recorded in scope §6.5: recovery discarding a run whose covered
identity was not preserved, which turns this refusal into reclamation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
The prerequisite for checkpointing. `recovery_generation` was two numbers
wearing one name: the logical generation of the segment recovery seals, and
the generation of the manifest recovery installs. Every index run publishes a
manifest, so every seal moved the number — and the segment recovery then
wrote took the moved value while the frames inside it were already named by
the old one. Persisted index locations dangled.
Split, each answers its own question. The logical generation is the tail's,
derived from the manifest's committed prefix, so sealing a journal into a
segment changes where the bytes are and not what they are called. The
manifest generation is the next free one, so an index run's manifest and a
recovery's cannot collide. Collision detection moved with them: the old check
compared against a maximum mixing all three namespaces, which an index run
could raise on its own, and only a segment can collide with a segment.
This closes the coverage restriction the previous commit had to impose. A run
may again cover locations naming the active tail, because the tail's identity
now survives being sealed away — so sealing covers the frames of the session
that wrote them instead of lagging one behind. Mutation-checked by putting the
segment's identity back on the manifest counter, which reproduces the original
defect exactly: a recovered run pointing at a logical generation nothing pins.
One case is not closed, and a test found it rather than review. An orphan
`.seg` from an interrupted seal occupies a logical generation whether or not
it is a readable segment, and it may hold exactly the identity the tail wants.
`an_orphan_segment_leaves_the_active_journal_the_authority` is also the test
documenting why recovery must not refuse there — the active journal is still
the authority and an outage would be the wrong answer — so the seal falls back
to a free generation and renames the frames, as it did before. An index run
against the old identity then dangles. Confined to roots carrying an orphan
segment, where it was previously universal; the closure is for recovery to
discard runs whose identity was not preserved, recorded in scope §6.5.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
Four findings against the previous commit. Each fix carries a regression that
fails against the landed code, and each was mutation-checked back to it.
Admission projected nothing. The replay-ceiling check read the published root,
so it decided about a transaction it had not counted: one object present, a
two-object transaction against a ceiling of two committed and the store then
failed to reopen. The open group was the same hole one step along. Admission
now projects sealed runs, unsealed layers, the open group, and the incoming
transaction.
The byte ceiling was unguarded. Recovery rebuilds into one delta that refuses
on either ceiling, so narrow frames across namespaces passed admission and
failed to reopen on `max_active_index_bytes`. Both are checked, and the
projection counts namespaces because the encoding pays a section header per
namespace. `index::encoded_bytes_for` is that arithmetic extracted, so this
file does not carry a copy of the encoding's shape.
A recovered run's locations did not resolve, and this reshaped the slice. An
`IndexLocation` names a logical generation, and a run is the first thing here
that persists one across a session — sound only for a generation that is
stable, which is a segment's alone. The active tail's is assigned from
`max(manifest, .seg, .idx) + 1`, so it moves whenever any artifact appears
(the run's own manifest suffices), and recovery seals a journal holding
frames at that counter rather than at the generation the tail had. The
previous reopen test could not see it: its lookups were answered by the
replay delta shadowing the run. Coverage is now an oldest-first prefix of
layers whose every entry is segment-backed, which makes the broken run
unwritable rather than untested. The cost — sealing lags one session behind
until frames leave `active/` — is recorded in scope §6.5.
Preserving the tail's generation across the seal was attempted and withdrawn.
`recovery_generation` is at once the new manifest's generation and the sealed
segment's logical generation, so the real fix separates those two numbers in
A2's recovery core, and that belongs with checkpointing rather than inside a
B1 integration commit. The first attempt also targeted the wrong branch: a
journal holding frames is replaced, not kept. `active_tail_logical_generation`
is left extracted at the one path that already used that formula so the two
ways of numbering an active tail are visible together.
The run ceilings counted every shard, where recovery enforces them against one
shard's manifest — a four-shard root with `max_index_runs = 1` refused the
second shard its first run. Counted per shard now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
A shard now seals its accumulated delta into an `IndexRun`, publishes it
through the manifest, and discards exactly the layers that run covers in one
committed-root CAS. This replaces the `NotImplemented` that refused a shard
once it had accumulated `max_index_runs` delta layers.
The seal runs at admission, before anything is reserved or sequenced, because
a seal that fails partway has to poison and poisoning a shard that has just
accepted a transaction owes that caller an answer it can no longer give. Two
triggers: `DeltaPressure::SealRequired` over the accumulated entries and
bytes, and the layer count, which is the fan-out every lookup pays before it
reaches a run and which entry pressure alone does not bound.
Everything up to `install_index_run` is pre-durable and fails as an ordinary
error; from there the shard poisons on any failure, including one the call
may have made before writing. The caller cannot distinguish those, and the
conservative direction is refusing to keep writing against a root that may no
longer describe the device. `poison_now` latches it — `poison_error` only
built the error, which is right inside the publication window whose caller
latches for the whole group, and wrong here.
Four frozen amendments, all recorded as contract review 2026-07-29-C:
`segment::install_index_run` holds the whole durability sequence so no
durability operation lives in `engine.rs`; `index::delta_pressure` becomes a
free function so the writer's multi-layer backlog asks the same watermark
rather than restating it; `CommittedRoot::merge` recognizes a publication
that appends no frame, without which the run reaches the manifest and never
the root; and a manifest may have an empty retained tail when it commits
through zero, which is every shard that has sealed an index but not yet
rotated its journal.
A finding, recorded in scope §6.5 rather than papered over. Sealing moves
entries out of the layers but no frame out of `active/`, and the committed
prefix advances only on a checkpoint or a rotation — so recovery still
replays everything into one ceiling-bounded delta. The writer therefore
refuses once the replayable set reaches `max_active_index_entries`, closing a
hole that pre-dates this change: the old layer cap never bounded the summed
entries behind it. The consequence is that an entry-pressure seal lands
exactly on that ceiling and the next admission is refused; only a fan-out
seal leaves the shard able to continue. Entry-pressure sealing becomes useful
when `checkpoint()` can advance the prefix.
Mutation-checked in both directions. Reverting the discard reports one run
beside a three-layer backlog where the test requires zero, so the frozen
`IndexMaintenanceSnapshot` is load-bearing. Reverting the `merge` amendment
reports zero runs beside a three-layer backlog while `CURRENT` names the run:
the silent divergence, arriving quietly.
`store-bench` is untouched and stays preliminary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
The last three hazards of contract review 2026-07-28-D, closed together
because they are one invariant: no name the store *invents* beneath a root
may be reached through a link or resolve to an object of the wrong type. The
flags and descriptor checks are the safety property, so the mechanics go in
the `sys.rs` funnel where the next author looking for how this crate opens
files will find them.
Refusal semantics differ per name, because the names mean different things.
`write_fenced` adopts and empties an existing regular `.tmp` — that is
residue from an interrupted attempt at this exact write, and reusing the name
is how a retry works — which is why the open and the truncation had to be
separated: no flag combination truncates only regular files, so the type
check needs the descriptor first and `O_TRUNC` cannot be in the open.
`read_format` keeps three answers apart: absent stays `Io(NotFound)` because
startup states 1 and 3 depend on it, non-regular is `UnrecognizedLayout`, and
a corrupt regular marker keeps its decode error. `initialize_root` leaves the
caller's root and ancestors alone, adopts existing directories beneath it, and
classifies every planned entry before creating any missing one so a refusal
cannot half-extend the tree it refused.
Directory fences now go to descriptors already validated rather than
re-resolving the name, which would hand the fence to whatever the name
resolves to now instead of what was checked. The fence sequence is otherwise
identical on purpose: `engine.rs` asserts the count exactly.
Each protection was reverted independently and the witnesses observed:
- `write_fenced` — a 4096-byte file outside the root truncated and
rewritten through a live link, a file created outside the root through a
dangling one, and a fifo at the name blocking the open for the full
ten-second deadline: an unbounded startup hang from one `mkfifo`.
- `read_format` — a foreign `FORMAT` read in full, its `shard_count` and
`root_uuid` returned as this root's, so every file in the tree would then
be validated against a marker the store never wrote. Same hang on the
read side.
- `initialize_root` — returned `Ok(FormatMarker)`, reporting a working
store with its shard tree built outside the root. The preflight has its
own witness: `shards/00/active` left behind by a refusal that named
`shards/00/segments`.
The witnesses sit on the `segment` entry points. `StoreEngine::open`'s
classifier refuses a redirected root before any of this is reached, so a test
entering that way passes whether or not the protection exists — and
`RecoverySession::open`, `drive.rs` and `store-bench` all arrive without it.
Disclosed: the device-node residual now covers `FORMAT.tmp` too — one
`O_NONBLOCK` open before the `fstat` refuses it, still gated behind `mknod`
privilege inside a configured root. `rename_noreplace` needed no change;
`RENAME_NOREPLACE` fails `EEXIST` on an occupied target whether or not it is
a link.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
Lead fix, folded in. The previous repair accepted any `#[cfg(test)]` item whose
first line ended in `{`, which is the same latch in another spelling: a
semicolon-terminated `static` or `const` can open a block initializer there and
close with `};`, so the exemption ran past it into the next function. Verified
by reverting the condition — the scanner returned no offender at all for a
`std::fs::write` in the function following a `LazyLock` initializer.
Only `mod` and `impl` are accepted as braced shapes, being the two the crate
actually uses. Everything else fails the guard by name rather than being
bounded by a brace that may not be the item's own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
Two findings against the previous commit, both upheld.
The funnel guard's test exemption could still latch. Ending it at the next
column-zero `}` is right for a braced item and wrong for every other shape:
after `#[cfg(test)] use crate::test_support;` the first such brace belongs to
the *next* function, so all of it went unscanned. The scanner now reads the
attributed item's shape — braced items are exempt to their closing brace,
semicolon-terminated items exempt only themselves, and any third shape,
including an item header rustfmt split across lines, fails the guard. A shape
it cannot bound is not a shape it may assume is harmless.
It is now a function over `&str` with synthetic tests, which is the more
important half. Mutating real sources only probes the shapes those sources
happen to contain: no file in this crate has a semicolon-terminated
`#[cfg(test)]` item followed by production code, so no mutation of a real
file could have produced this defect. Charter item 8's analogue for tooling.
The two-owner claim was overstated. The regression arranges its wrong-typed
name by replacing `LOCK` under a live holder — and replacing it with a fresh
*regular* file succeeds just as well, since both opens are then of a regular
file at the right name with nothing to tell them apart. The type check closes
"the name already resolves to the wrong kind of object", the operator-error
and stale-state case; it does not close "the name is replaced under a
holder", and no check at this layer can.
So scope 3.1 now separates the two, says which is in scope, and states the
replacement case as an explicit deployment assumption rather than leaving it
implied: anything able to replace `LOCK` can equally unlink a journal, so
advisory locking was never the boundary that would stop it. The assumption is
pinned by a test asserting the current behavior on purpose — if a stable
locking object is ever adopted, that test is meant to fail, and the failure
is the signal that the documented assumption changed. §3.1 records locking
the root directory as the candidate and what it would cost.
Also exact rather than caveated: a Unix socket fails `open(2)` with `ENXIO`
before any `fstat`, so it surfaced as `Io` while the documentation promised
`UnrecognizedLayout`. `ENXIO` and `EISDIR` both now mean "not a regular
file", and the socket is one of four occupants the test loop covers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
`flock` locks the inode a descriptor reached, not the name that was asked
for, and `lock_root` opened `<root>/LOCK` with `create(true)` — a
follow-through open. A symlink at that name therefore put the root lock on
a foreign inode and left the root's own lock file unlocked, so a second
process arriving at the same root took the lock as well: two owners, each
believing it held one root exclusively. Nothing is destroyed and everything
downstream is permitted to race, which is why this is scheduled on its own
rather than folded into the remaining symlink work.
`classify_root` already refused a non-regular `LOCK`, but only on the
`StoreEngine::open` path. `RecoverySession::open` and `drive.rs` reach
`lock_root` directly with no classification ahead of them, so the guarantee
had to move down to the open itself.
`sys.rs` gains a third no-follow primitive for the shape the other two do
not cover — a name the store must own and keep, adopting an existing regular
file and creating an absent one, never truncating. The type is established
by `fstat` on the descriptor already held, before `flock` is attempted, so a
refused name is never locked even momentarily. A non-regular occupant is
`UnrecognizedLayout`, not `AlreadyLocked`: the root is malformed, not busy.
Measured on a reverted copy, three distinct failures rather than one:
- two owners of one root, with the first lock still held;
- a dangling link at `LOCK` created a file outside the root;
- a fifo at `LOCK` returned `Ok(RootLock)`, the store reporting that it
held the root lock on a pipe. That one was found by writing the test
for the type check, not predicted.
The funnel guard needed amending to accept these tests, and the reason it
did is a defect in the guard: it exempted test code by matching the literal
name `mod tests`, so the two modules named otherwise were scanned as
production code while a file could have evaded the guard entirely by naming
a module `tests`. It now keys on the `#[cfg(test)]` attribute and, unlike
before, the exemption ends at the module's closing brace — code appended
after a test module used to be unscanned. Both directions mutation-checked.
Disclosed, not closed: a device node at `LOCK` still receives one
`O_NONBLOCK` open before the `fstat` refuses it. The three remaining
symlink hazards in `segment.rs` stand unfixed; contract review 2026-07-29-A
records both.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKzM69CHmBuDcA3qN1jFdh
A bundle's verification flags were assertions about methodology that
nothing checked. This makes them conditional on machine-readable
declarations of what the run actually did, and fixes five cases where the
emitter stated something it had not established.
bench/result-schema.json gains a required run_conditions block of ten
closed enumerations -- initialization and mutation path, checkpoint and
index state, index-run ceiling, receipt reconciliation, objects_new
source, commit-id uniqueness, build profile, environment fidelity. It is
not disclosure beside the claims; it is what the claims are conditioned
on, so a harness can only assert what its declaration permits. A prose
caveat field was rejected: free text is not a condition a consumer can
check, and a bundle whose caveats live only in a report reads as
unconditional to everyone who receives it.
The branch conditional forbids the three newly earnable claims on the
journal-drive path and pins its four provenance declarations to the only
values that seam can make. Forbidding the claims alone left the hole one
field over -- a drive bundle could otherwise declare exact receipt
reconciliation it has no receipts to perform.
Emitter defects, each found by reading the schema against the code:
- Setup traffic was inside the measured interval. Both counter baselines
were read only at the end, so repository creation -- which goes through
submit, and therefore fences and signs -- was counted as measured work
while the bundle asserted setup_traffic_excluded. A false exclusion
claim is worse than a wrong number: a wrong number invites scrutiny
and this deflects it.
- A zero-work run produced a schema-valid bundle asserting uniqueness
over zero ids and three-objects-per-commit over zero commits. Both are
vacuously true, which is why they must not be earnable that way: the
result is indistinguishable from a measured run by the consumer the
schema exists to serve. Refused by name at two altitudes.
- An ACK-journal write failure ended the run quietly. It set a stop flag
without recording a refusal, so neither the fatal guard nor the
zero-work guard saw it, and the bundle omitted a committed transaction
while still counting its fence and its signature -- one counted
transaction against two fences and two signings. It is now a fatal
incomplete-accounting refusal carrying the original errno, because the
commit happened: folding it into the refused count would report a
transaction the store committed as one it declined.
- Widening that class to "a failure that produces a value nobody read"
found three more. A shard with no counters summed to zero fences,
silently shrinking the total that bounds every durability claim. A
digest of an unreadable file returned the digest of empty input -- a
well-formed 64-hex value indistinguishable from a real one, feeding
five attested fields. An unreadable /proc/meminfo published one byte
of RAM. All three refuse now.
- Index steady state was inferred from any directory entry, so one stray
file declared the index sealed. Entries are parsed back as index runs
against the root's own uuid; an unparseable entry is reported as
unvalidatable rather than lowering a count, and a backlog is refused
because neither named value describes sealing that did not keep up.
deployment.tmpfs, persistent_data_mount, and hardware.filesystem were
constants -- the emitter could assert deployment facts it had never
checked. They are read from /proc/mounts now. Both deployment fields relax
to booleans so a diagnostic run is representable at all: it was previously
not disqualified but unencodable, and a schema that can only express
successful runs is not a record of what was measured. outcome=pass
requires reference fidelity at every gate, and a diagnostic run may never
carry a pass verdict.
Hardware profiles are derived, never accepted. store-bench parses the
whole frozen profile tables and names a profile only by exact comparison,
iterating every pinned fact rather than every supplied one -- so a fact
the emitter does not model eliminates the profile instead of being
invisible. That took the honest unobserved list on this host from 7 facts
to 24, which is the inversion working. A deployed-node harness supplies
privileged facts as evidence to compare, never as a label. The outcome
derivation now also requires the checkpoint and index conditions, because
fidelity was the only thing preventing a pass and would have stopped being
so the moment profile recognition started working, at which point the
emitter would have produced a pass its own validator rejects.
operation_receipts_reconciled is expressible and deliberately not emitted:
the bench accepts any Committed status without comparing the payload, and
the digest it records is of the operation id rather than the receipt.
Contract review 2026-07-28-C records the amendment and the ceiling it does
not close: run_conditions is self-reported, and only the index-run ceiling
is cross-checked against an independent value. Scope 5.1 records why a
full filesystem is indistinguishable from a concurrency flake by symptom,
and that an I/O error must reach a report with its errno intact -- the
same requirement as the incomplete-accounting refusal above.
scripts/check-phase1.sh GATE_EXIT=0; verify-store-recovery.sh reports
bundle=schema-valid on both paths, zero_work_run=refused, and
unaccounted_ack_run=refused.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scope 6.4 deliverable 1, startup state 1. StoreEngine::open now creates a
root rather than refusing to, and every harness that measured a store
production did not build is re-pointed at it. States 3 and 4 still refuse.
The four startup states are distinguished by one read-only classification
that creates nothing. A FORMAT entry means state 2 on the strength of the
name alone, because a FORMAT that does not decode is still FORMAT and
treating an unreadable marker as "no marker, therefore empty" would
authorize building a fresh tree over a populated root. A root holding only
LOCK is empty, since lock_root creates that file as a side effect of asking
whether the root is busy. The signer check moved above every root access,
because once an absent root is initialized rather than refused, a signerless
configuration would otherwise create a tree, write FORMAT and fsync the
parent before failing -- a configuration error must not leave a root behind.
Three defects found in review, all fixed here rather than deferred.
A deeply absent path was created with create_dir_all and only its immediate
parent fenced, so open could report success over ancestors a power loss
could take. Ancestors are now created one at a time -- so what this call
created is exactly what it fences, and a racing creator surfaces as
AlreadyExists rather than being absorbed -- and fenced deepest-first, since
a directory entry lives in the parent that names it and the reverse order
can leave a fenced parent naming an unfenced child.
An interrupted initialization was unrecoverable: any tree residue classified
the root as non-empty-without-FORMAT and it was refused forever, with a
message about legacy layouts that had nothing to do with what happened. A
root being built now carries an INITIALIZING marker, installed by rename as
the first durable act and removed as the last, so a durable partial tree
always has a durable marker beside it and every crash window is resumable.
Recognizing it requires both a byte-exact marker this crate alone writes and
every entry in the root drawn from a closed set of names this store
invented, so a foreign layout cannot be mistaken for abandoned
initialization and overwritten -- the direction that matters, since refusing
a resumable root costs an operator time and overwriting a real one costs
their data. Sibling staging with atomic installation was the alternative and
is structurally blocked: LOCK lives inside the root, so the root must exist
before any mutation can be serialized, and renaming a tree onto a directory
containing LOCK fails ENOTEMPTY.
The classifier then ignored INITIALIZING.tmp by name regardless of type or
contents, and initialization opened that name with create plus truncate. An
operator's file there was destroyed silently, and a symlink there truncated
a file outside the root to 27 bytes and then removed the link -- destroying
data the store never owned and erasing the evidence, while open returned Ok
and reported a working store. The justification for ignoring the name was
that only this path could have written it, which is circular: that is the
claim the classifier runs in order to establish. Every entry is now judged
by lstat type before anything opens it, the temporary marker is validated as
an exact regular marker or refused, and installation is create-new rather
than create-truncate. Contract review 2026-07-28-D records the two no-follow
open primitives this added to the frozen sys.rs, and the four further
symlink hazards in segment.rs that are recorded rather than fixed -- the
first of which lets two processes believe they hold one root lock.
A fifo at that name made the pre-fix open block forever: one mkfifo in a
configured root was an unbounded startup hang, not only a data hazard.
The engine and the drive seam are now asserted to recover one crash image
identically, closing a gap that was true by construction and untested.
Charter item 8 applied to the harness: the in-crate test helper no longer
calls segment::initialize_root, so every writer test builds its root through
open; the ROOT_SEEDED_BY_NON_PRODUCTION_PATH disclosure is retired; and the
fixture's root_seeded_by becomes a stable token matched by exact equality,
with the history moved to an adjacent reason field -- a substring match
passes on a value that has drifted to mean something else.
The d0 contract test asserting open returns NotImplemented for any valid
configuration is obsoleted by this deliverable and replaced with the
stronger property: a signerless configuration is refused and leaves no root
behind. It moves off a fixed /tmp path, which under the old check ordering
would have created a real store root on every gate run on every machine.
scripts/check-phase1.sh GATE_EXIT=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
B4's harness measured intermittent AlreadyLocked on reopening a store it
had just dropped -- up to 29 retries over 147ms, 7 failures in 40 runs,
reproducing on a single-shard engine. StoreEngine::drop was not the cause:
it closes every channel and joins every writer, and looked correct in
every single-process test.
An flock is held by the open file description, not by a descriptor. A
concurrently forked child transiently inherits that description, and
FD_CLOEXEC closes it at exec, not at fork -- so during that window the
parent closing its own descriptor releases nothing. Scope 3.1 makes
AlreadyLocked a refusal and never a wait, so every consumer that closes
and reopens a root -- a recovery drill, an in-place restart, the Phase 2
migrator -- could be refused with no defined retry.
Releasing must therefore be an explicit act. lock_root returns an RAII
RootLock that issues LOCK_UN in Drop before the file closes, and both
owners -- RecoverySession and ShardDrive -- hold that guard. The unlock
lives in sys.rs beside try_lock_exclusive so the locking syscalls stay in
the one funnel. StoreEngine needs no change: it holds a RecoverySession,
never a File, so it inherits the fix.
Drop is the only release path. A release() and a file() accessor were
written and deleted before landing: neither had a caller, and an uncalled
second way to release a lock is exactly the decoy the charter names. A
failed LOCK_UN in Drop cannot be returned and must not be swallowed, so it
increments a counter the regression asserts unchanged.
The rejected alternative was fixing this in StoreEngine::drop alone. That
leaves ShardDrive and the drive's one-shot session exposed and makes
correctness depend on a descriptor lifetime that fork can extend.
The regression is synchronized rather than timed: the child forks while
the lock is held, signals ready on one pipe, and blocks on a second until
after the parent has released and attempted its reopen, so the inherited
descriptor is provably open across the whole window and the reopen is
asserted on its first attempt. With the explicit unlock removed it fails
10/10; as landed it passes 40/40. Measured under load -- 200 close-reopen
cycles against 72,255 concurrent forks -- 0 refusals, worst case 1 attempt
and 10.8ms; the same load kills the pre-fix behaviour within 0.05s, so the
load reproduces the defect rather than merely being weak.
Contract review 2026-07-28-B records the amendment. Two carry-forwards are
recorded in scope 6.6: the verify-store-recovery SIGKILL cycles still drive
the journal seam rather than submit, so kill -9 never lands inside a real
publication and the acknowledged-crash-recovery criterion is only partly
earned; and B4's bounded reopen retry must become a one-attempt assertion
now that the defect it compensates for is gone.
scripts/check-phase1.sh GATE_EXIT=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The B2 review of the first B1/B3 slice found five defects whose fixes were
not available to the packages that had to make them: each needed a change
to a surface those packages do not own. Rather than let them restate a
frozen fact locally, the lead amends the surfaces and the packages consume
them. Contract review 2026-07-28-A records all five with the weaker
alternative that was rejected for each.
- CommitEvidenceSigner::sign_event's parameter becomes signing_digest.
The message is SignedCommittedTransactionV1::signing_digest, which
binds key epoch and durability result around the canonical event; the
bare event digest is chain identity, not signature material. A doc line
was not enough: misuse is undetectable until after durability, and is
first observed by a mirror on another instance.
- TransactionEvidenceV1::actor() is new, exhaustive over all six
variants. A destination event's actor must restate the evidence's
source instance, and deriving it a second way store-side is what
produced the defect it closes.
- format::object_type_code becomes pub(crate) and recovery.rs's private
twin is deleted, so one table exists where three did.
- RefRecord gains from_target/target() in terms of RefTarget, with the
code table stated once per direction and an unknown kind refused by
value as CheckpointError::RefKind. Defaulting an unknown kind would
launder it into the next checkpoint within one interval.
- max_projection_objects is capped at MAX_CANONICAL_ITEMS and its default
lowered to it; max_projection_chunks likewise, and max_projection_bytes
against the transitive per-chunk allowance. The old default described a
projection no manifest could encode. Supporting a hundred million
objects requires a versioned chunked or indexed manifest design, not a
larger hostile-decode ceiling.
Both governing documents are updated where they now misstate a frozen
fact, including §6.4's claim that signing covers the event digest. §6.5
records the two carry-forwards from B3's first slice — production staging
use is incomplete, and expire() has no scheduler — and records that the
requested StoreEngine staging accessor is a pending amendment which must
not be a bare Arc<ProjectionStaging>, since a clone could outlive
EngineShared, survive release of the root LOCK, and keep serving a root
this process no longer holds.
No package file is touched: engine.rs, transaction.rs, and staging.rs
consume these in their own commits.
levcs-protocol 6 consumer tests, levcs-store 109 library tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The freeze commit cannot contain its own hash, so it is recorded here.
Wave A is frozen at 5ee9c6b78b with the gate
evidence measured at that commit, in both the plan and the scope document.
Restates what the freeze does and does not cover: the two carry-forwards are
B-wave work and outside it, and doc/swarm-fabric-roadmap-exploration.md
(e6a058d) stays non-binding and outside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG
Deliberately a separate commit, outside the Wave A freeze (5ee9c6b). This
document decides nothing and is not part of any frozen contract: it is a place
to hold the thinking on whether LeVCS should become a change fabric for
multi-agent work, pending its own review.
Records where the proposal maps onto the existing plan, where it cuts against
it -- levcsd must not own ref transactions, the ephemeral stratum invariant,
delegation being identity-shaped rather than protocol-shaped, attestations not
fitting TransactionEvidenceV1, and the canonical benchmark workload being the
wrong shape for swarm traffic -- and where the proposal is simply wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG
Freezes the Wave A interfaces, frame format, durability ordering, and
crash/fault fixtures per doc/phase1-storage-spine-scope.md section 5. Wave B
may now start against these.
D0 lands crates/levcs-store: the sealed public API, the file-ownership split,
the single durability syscall funnel with its counters and fault hooks, the
17-entry failpoint registry in compiler-enforced correspondence with
oracle::AppendFailpoint, and the journal-level drive seam. Wave A lands the
frame codec and journal/segment lifecycle (A1), the recovery index and
checkpoints (A2), and the durable-ingest benchmark and crash harness (A3).
The adversarial review found five defects behind a green gate, three of them
blockers, all closed here. Two were the same shape: drive::reopen_through_
recovery had reimplemented a simplified recovery and called none of recovery.rs
-- so it adopted segment footer sequences without validating frame bytes,
journal_id, or root_uuid, and it double-adopted interrupted seals. The seam
between two packages was untested precisely because each package's own tests
passed. reopen_through_recovery now delegates rather than decides, and
DriveRecovery carries the recovery report verbatim so tests assert the
disposition and not merely its effect: a double adoption and a correct replay
produce the same adopted set, which is how the defect stayed invisible.
Also closed: store-bench now schema-validates its own emitted artifact with a
five-mutation negative control instead of matching JSON substrings; the ACK
reconciler distinguishes duplicates and regressions from forward gaps;
checkpoint writes and a journal truncation are routed through the durability
funnel, whose guard now covers writes and truncations rather than only sync,
rename, and unlink.
bench/result-schema.json is amended (contract review 2026-07-24-B, second and
third amendments): per-gate latency ceilings conditional on outcome so a failed
run is representable, and the verification claims split per gate so a storage
run cannot certify an object graph it never touches. Not-applicable claims are
forbidden rather than falsified; applicable-but-not-performed report false.
Every relaxation is re-pinned in the else branch and asserted member by member,
after an edit in this series silently un-pinned all eleven validation flags and
was caught only by revalidating against constructed bundles.
Arming the fault registry now requires a FaultSerial token, so the invariant is
a compile error rather than a comment. The file where this was diagnosed
carried a header saying it was deliberately the only test in it, and a second
test had been added under that comment anyway -- an 8-in-40 failure rate that
read as flakiness.
Evidence at this commit: check-phase1.sh GATE_EXIT=0 across all four feature
configurations, 124 test binaries, zero failures; verify-store-recovery.sh
--cycles 100 with recovery_failures=0, acknowledged_loss=0,
torn_transactions=0, repeated_adoptions=0, bundle=schema-valid; recovery_eio
40/40 at four test threads; golden corpus byte-stable; fmt clean.
Carry-forwards, explicitly not Wave A blockers and recorded in scope section 5:
extend the crash matrix to generate sealed-frame corruption and cross-shard
journal movement, since it structurally cannot express the class the first
blocker belonged to; and wire GroupBuilder through B1's production path, since
deliverable 4-A1.2 is presently asserted only over a type nothing calls.
Charter item 9 -- ask every package what of its work is correct but uncalled --
is accepted for every subsequent wave. It, not the review, found the class.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QNy4Ve7mogg4X1ezJTnFxG