Two findings, both about a claim that could not be falsified.
The --self-test plan put the failing step last. With the failure last,
a runner that ABORTS on failure and one that CONTINUES produce
identical output, so the witness for Q#GR-2 policy --- the suite keeps
going --- would have passed on a runner doing the exact opposite. The
plan is now three lines with a passing SENTINEL after build-crdt,
asserted to have written its own log. That is the only thing that
distinguishes the two behaviours, and it turns Q#GR-2 from a declared
policy into an observed one.
The plan test also now pins the EXACT command, not only the step name
and its position. A build-crdt running plain cargo build would leave
the gate exactly as unsound while looking repaired --- the crdt sweep
needs those specific features, which is the whole defect.
The ledger still recorded the superseded boundary decision: "section 3
gains it, section 5 keeps the incident, and the script cites both".
Revision 2 replaced that with section 3 as the sole normative home and
the script citing section 3 alone. active-work.md is the volatile
cross-machine record, so a recovering machine reading the stale entry
would have rebuilt revision 1 wrong boundary. Now updated, and it says
which decision it supersedes rather than silently replacing it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Three review findings.
The normative build requirement goes entirely into handoff section 3.
Revision 1 proposed section 3 gaining it while the script header cited
both sections, which splits one executable contract across two homes
and weakens the single clean boundary the script has --- at the same
time as Q#GR-4 declines to build any automated check for prose drift. A
boundary that is neither enforced nor singular is not a boundary.
Section 5 keeps the incident and its signature, which is history rather
than contract.
Q#GR-1 observation procedure was unsafe and insufficient. "Delete
pmacs-gpu from a target directory" mutates a live worktree build
directory, and removing one binary does not establish that the other
artifacts and feature permutations are cold --- a stale dependency
graph can satisfy the run for reasons the experiment never sees. Now: a
disposable target, the binary asserted ABSENT before each run as a
recorded precondition, and the two sweeps run separately so neither can
be explained by the other having built the binary first. That last
point is the same accident that hid this defect for the whole life of
the shared target dir.
The attribution criterion had no feasible witness. gate_script_acceptance
deliberately runs no gates, so plan assertions prove name and order and
nothing about runtime behaviour. The obvious seam is a trap: making
PLAN_FILE injectable would turn the script into a general command
executor through its runner eval --- the same class of defect this
script own review already caught in --acceptance and fixed with a
parse-time refusal. Reintroducing it one lane later, in the tool whose
purpose is to be trustworthy, is not a trade worth making.
Q#GR-5 proposes --self-test over a HARDCODED two-line synthetic plan,
true and false, with the failing one named build-crdt. No injection,
no real gate, and it tests the thing actually under test: whether the
runner names the right gate when a command fails. Whether cargo build
really fails is cargo business.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
--protocol promises the CRDT workspace sweep. That sweep documented
precondition is cargo build --workspace --no-default-features
--features luajit,crdt (handoff section 5:532-535), and the plan
emitter at scripts/gate:187-204 has no build step at all --- read from
the source, not inferred from the failure.
The interesting part is why it stayed invisible. Before #225 every
worktree on this machine resolved to one shared CARGO_TARGET_DIR, which
almost always already contained a pmacs-gpu binary, so the precondition
was satisfied by accident on essentially every run. Per-worktree target
dirs start empty. So this is not a bug #225 introduced; it is a
pre-existing gap in the documented procedure that #225 stopped hiding.
That also decides the urgency. A red gate is fine --- it stops you. The
hazard is the reverse: a GREEN --protocol run whose crdt sweep was
decided by what happened to be in the build directory rather than by
the diff. A gate reporting coverage it does not have is exactly what
#225 exists to prevent, so the tool shipping with this gap teaches the
opposite of what it is for.
Observed on PR #228 first gate run: twelve
gpu_invocation_acceptance::crdt::* failures, all "build pmacs-gpu
before this acceptance suite", with debug/pmacs-gpu absent from the
fresh target dir.
The durable half is a boundary question rather than a missing line. The
script header names handoff section 3 as the owner of its reasoning,
and this precondition lives in section 5 --- a coherent cause for the
omission, not oversight. Q#GR-3 proposes section 3 gains it, section 5
keeps the incident and its signature, and the script stops naming
section 3 as its only source.
Q#GR-1 is marked as the one thing this lane will not accept on
reasoning: whether the default sweep also needs the binary must be
established by deleting it and running both sweeps. The whole defect is
a precondition nobody checked, and establishing its replacement by
reading would repeat the error at one remove. The mechanism section
states its own inference (the failing tests are namespaced ::crdt:: and
so are probably feature-gated) and marks it unverified.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai