LeVCS/bench
Levi Neuwirth 062797d4b8 Condition benchmark claims on the run that produced them
A bundle's verification flags were assertions about methodology that
nothing checked. This makes them conditional on machine-readable
declarations of what the run actually did, and fixes five cases where the
emitter stated something it had not established.

bench/result-schema.json gains a required run_conditions block of ten
closed enumerations -- initialization and mutation path, checkpoint and
index state, index-run ceiling, receipt reconciliation, objects_new
source, commit-id uniqueness, build profile, environment fidelity. It is
not disclosure beside the claims; it is what the claims are conditioned
on, so a harness can only assert what its declaration permits. A prose
caveat field was rejected: free text is not a condition a consumer can
check, and a bundle whose caveats live only in a report reads as
unconditional to everyone who receives it.

The branch conditional forbids the three newly earnable claims on the
journal-drive path and pins its four provenance declarations to the only
values that seam can make. Forbidding the claims alone left the hole one
field over -- a drive bundle could otherwise declare exact receipt
reconciliation it has no receipts to perform.

Emitter defects, each found by reading the schema against the code:

  - Setup traffic was inside the measured interval. Both counter baselines
    were read only at the end, so repository creation -- which goes through
    submit, and therefore fences and signs -- was counted as measured work
    while the bundle asserted setup_traffic_excluded. A false exclusion
    claim is worse than a wrong number: a wrong number invites scrutiny
    and this deflects it.
  - A zero-work run produced a schema-valid bundle asserting uniqueness
    over zero ids and three-objects-per-commit over zero commits. Both are
    vacuously true, which is why they must not be earnable that way: the
    result is indistinguishable from a measured run by the consumer the
    schema exists to serve. Refused by name at two altitudes.
  - An ACK-journal write failure ended the run quietly. It set a stop flag
    without recording a refusal, so neither the fatal guard nor the
    zero-work guard saw it, and the bundle omitted a committed transaction
    while still counting its fence and its signature -- one counted
    transaction against two fences and two signings. It is now a fatal
    incomplete-accounting refusal carrying the original errno, because the
    commit happened: folding it into the refused count would report a
    transaction the store committed as one it declined.
  - Widening that class to "a failure that produces a value nobody read"
    found three more. A shard with no counters summed to zero fences,
    silently shrinking the total that bounds every durability claim. A
    digest of an unreadable file returned the digest of empty input -- a
    well-formed 64-hex value indistinguishable from a real one, feeding
    five attested fields. An unreadable /proc/meminfo published one byte
    of RAM. All three refuse now.
  - Index steady state was inferred from any directory entry, so one stray
    file declared the index sealed. Entries are parsed back as index runs
    against the root's own uuid; an unparseable entry is reported as
    unvalidatable rather than lowering a count, and a backlog is refused
    because neither named value describes sealing that did not keep up.

deployment.tmpfs, persistent_data_mount, and hardware.filesystem were
constants -- the emitter could assert deployment facts it had never
checked. They are read from /proc/mounts now. Both deployment fields relax
to booleans so a diagnostic run is representable at all: it was previously
not disqualified but unencodable, and a schema that can only express
successful runs is not a record of what was measured. outcome=pass
requires reference fidelity at every gate, and a diagnostic run may never
carry a pass verdict.

Hardware profiles are derived, never accepted. store-bench parses the
whole frozen profile tables and names a profile only by exact comparison,
iterating every pinned fact rather than every supplied one -- so a fact
the emitter does not model eliminates the profile instead of being
invisible. That took the honest unobserved list on this host from 7 facts
to 24, which is the inversion working. A deployed-node harness supplies
privileged facts as evidence to compare, never as a label. The outcome
derivation now also requires the checkpoint and index conditions, because
fidelity was the only thing preventing a pass and would have stopped being
so the moment profile recognition started working, at which point the
emitter would have produced a pass its own validator rejects.

operation_receipts_reconciled is expressible and deliberately not emitted:
the bench accepts any Committed status without comparing the payload, and
the digest it records is of the operation id rather than the receipt.

Contract review 2026-07-28-C records the amendment and the ceiling it does
not close: run_conditions is self-reported, and only the index-run ceiling
is cross-checked against an independent value. Scope 5.1 records why a
full filesystem is indistinguishable from a concurrency flake by symptom,
and that an I/O error must reach a report with its errno intact -- the
same requirement as the incomplete-accounting refusal above.

scripts/check-phase1.sh GATE_EXIT=0; verify-store-recovery.sh reports
bundle=schema-valid on both paths, zero_work_run=refused, and
unaccounted_ack_run=refused.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 17:38:03 -04:00
..
workloads Freeze Wave A: Phase 1 storage spine 2026-07-26 19:47:03 -04:00
reference-hardware.toml Freeze Wave A: Phase 1 storage spine 2026-07-26 19:47:03 -04:00
result-schema.json Condition benchmark claims on the run that produced them 2026-07-29 17:38:03 -04:00