fix(gate): build pmacs-gpu before the crdt sweep, and witness the runner

`scripts/gate --protocol` emitted `sweep-crdt` with no build step. The
crdt sweep spawns `pmacs-gpu` as a process, and nothing in a
`cargo test` run produces that binary --- `pmacs-gpu` has no `tests/`
directory, so cargo never uplifts its bin to `debug/pmacs-gpu`. On a
cold target directory the sweep therefore fails twelve
`gpu_invocation_acceptance::crdt::*` tests on "build pmacs-gpu before
this acceptance suite".

The hazard was never the red gate. Before per-worktree target
directories (#225) every worktree shared one, which nearly always
already held the binary, so the precondition was satisfied BY ACCIDENT
for the whole life of that arrangement --- a GREEN `--protocol` run
whose crdt sweep was decided by the state of the build directory rather
than by the diff.

Q#GR-1 SETTLED BY OBSERVATION, not by reading. On a disposable target
directory with `debug/pmacs-gpu` asserted ABSENT before each run
(recorded, not assumed), each sweep run alone from the same cold state:

  default  cargo test --workspace --no-fail-fast -- --skip basedpyright
           exit 0, 114 test targets green, and `debug/pmacs-gpu` was
           STILL ABSENT afterwards --- the default sweep never builds it
           and never needs it.
  crdt     cargo test --workspace --features crdt --no-fail-fast
           -- --skip basedpyright
           exit 101, exactly twelve failures, all
           `gpu_invocation_acceptance::crdt::*`, matching the signature
           handoff section 5 recorded.

So the step is conditional on `--protocol`, which the framing voted for
on an inference this run confirms rather than assumes.

Also observed, and worse than the twelve: `a54_real_daemon_real_pty_and_
headless_gpu_render_one_panel_hosted_terminal` reported `ok` in that
same cold crdt sweep. Its only path that does not spawn `pmacs-gpu` is
its skip branch, so a test whose whole purpose is real wgpu rendering
passed having rendered nothing. The missing build does not only fail
twelve tests --- it silently voids coverage in tests that report green.

A NAMED STEP, NOT A FOLDED COMMAND. `cargo build ... && cargo test ...`
would report a BUILD failure under the name `sweep-crdt`, a wrong
attribution in the one place this script exists to be trustworthy
about.

`--self-test` is how that attribution is witnessed at all. The existing
suite drives only no-gates paths, so plan assertions can prove a step's
name and order and NOTHING about what the runner does when a step
fails. The mode runs a HARDCODED three-line synthetic plan through the
real runner: a passing step, a failing one named `build-crdt`, and a
passing SENTINEL after it. The sentinel is load-bearing --- with the
failure last, an aborting runner and a continuing one produce identical
output, so the witness would pass on a runner doing the opposite of the
stated policy.

The plan is a literal inside the script. Making `PLAN_FILE` injectable
would work and would turn the runner's `eval` into a general command
executor --- the same defect this script's own review caught in
`--acceptance` and fixed with a refusal at parse time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
This commit is contained in:
Levi Neuwirth 2026-08-09 16:39:14 +02:00
parent 12affd78e1
commit f55ce54627
No known key found for this signature in database
2 changed files with 248 additions and 4 deletions

View File

@ -6,8 +6,11 @@
# scripts/gate --print-target-dir
# scripts/gate --init
# scripts/gate --prune [--force]
# scripts/gate --self-test
#
# Framing: docs/gate-script-framing.md (revision 4, approved).
# Framing: docs/gate-script-framing.md (revision 4, approved), and
# docs/gate-protocol-build-framing.md (revision 3, approved) for the
# crdt build step and --self-test.
#
# WHY A PER-WORKTREE TARGET DIRECTORY. This machine exports one
# CARGO_TARGET_DIR for every checkout, and cargo takes an EXCLUSIVE LOCK
@ -40,6 +43,7 @@ usage: scripts/gate [--acceptance SUITE]... [--protocol] [--print-plan]
scripts/gate --print-target-dir
scripts/gate --init
scripts/gate --prune [--force]
scripts/gate --self-test
--acceptance SUITE a touched acceptance suite to run (repeatable).
docs/agent-handoff.md section 3 stays authoritative
@ -48,7 +52,8 @@ usage: scripts/gate [--acceptance SUITE]... [--protocol] [--print-plan]
working tree, and one that guessed would report
coverage it does not have.
--protocol the change touches PROTOCOL_VERSION; adds the CRDT
workspace sweep on top of the default one.
workspace sweep on top of the default one, plus
the build that sweep needs (see build-crdt below).
--print-plan print the exact gate commands and exit.
--print-target-dir print this worktree's build directory and exit.
Creates nothing.
@ -57,6 +62,13 @@ usage: scripts/gate [--acceptance SUITE]... [--protocol] [--print-plan]
--prune list managed directories whose worktree is gone.
Deletes NOTHING without --force.
--force with --prune, actually delete.
--self-test drive the real runner with a HARDCODED synthetic
plan --- true, false, true --- to witness that a
failing gate is named as ITSELF and that the suite
CONTINUES past it. Runs no real gates. EXITS
NON-ZERO BY DESIGN: the middle step fails on
purpose, so a non-zero status is this mode
working, not this mode broken.
EOF
exit 2
}
@ -198,11 +210,80 @@ emit_plan() {
if [ "$PROTOCOL" = 1 ]; then
# Section 3: touching PROTOCOL_VERSION STRENGTHENS the sweep
# line, it does not replace it. Both sweeps run.
#
# THE BUILD IS A PRECONDITION OF THE SWEEP, not a courtesy. The
# crdt sweep spawns `pmacs-gpu` as a PROCESS, and no `cargo
# test` run produces that binary: pmacs-gpu has no tests/
# directory, so cargo never uplifts its bin to debug/pmacs-gpu.
# On a cold target directory the sweep therefore fails twelve
# gpu_invocation_acceptance::crdt::* tests on "build pmacs-gpu
# before this acceptance suite" --- and, worse, other crdt tests
# that render through the real binary SKIP THEMSELVES and report
# ok, so the missing build also voids coverage silently.
#
# WHY ONLY UNDER --protocol, measured rather than reasoned. On
# 2026-08-09, on a disposable target directory with
# debug/pmacs-gpu asserted ABSENT before each run and each sweep
# run alone from that cold state: the DEFAULT sweep exited 0 and
# left debug/pmacs-gpu still absent --- it never builds the
# binary and never needs it --- while the crdt sweep exited 101
# with exactly those twelve failures. So the default gate does
# not pay for this build.
#
# A SEPARATE NAMED STEP, never folded into the sweep command.
# `cargo build ... && cargo test ...` would report a BUILD
# failure under the name `sweep-crdt`, which is a wrong
# attribution in the one place this script exists to be
# trustworthy about. --self-test is what witnesses that the
# runner names the failing gate as itself.
printf 'build-crdt\tcargo build --workspace --no-default-features --features luajit,crdt\n'
printf 'sweep-crdt\tcargo test --workspace --features crdt --no-fail-fast -- --skip basedpyright\n'
fi
printf 'diff-check\tgit diff --check\n'
}
# ---------------------------------------------------------------------
# The synthetic plan behind --self-test: the runner held to its own
# contract.
#
# WHY A MODE EXISTS AT ALL. tests/gate_script_acceptance.rs drives only
# NO-GATES paths --- a test that ran the real suite would run the gate
# suite inside the gate suite --- so plan assertions can prove a step's
# name and its order and NOTHING about what the runner does when a step
# fails. That left two stated properties with no way to observe them:
# a failing gate is attributed to ITSELF (which is the whole reason
# build-crdt is a separate step rather than `cargo build && cargo
# test`), and the suite CONTINUES past it rather than aborting. Same
# shape of argument as --init: verification needs a path it can drive
# safely, and this one is not a second implementation --- it hands the
# REAL runner loop a different plan file.
#
# THE PLAN IS A LITERAL, and that is the design, not a shortcut. The
# obvious seam --- letting a caller supply PLAN_FILE --- would work,
# and it would turn the runner's `eval` into a general command
# executor. That is the same class of defect this script's own review
# caught in --acceptance and fixed with a refusal at parse time;
# reintroducing it in the tool whose purpose is to be trustworthy is
# not a trade worth making. Nothing external supplies a command here.
#
# THREE LINES, AND THE THIRD IS LOAD-BEARING. With the failure LAST, a
# runner that aborts and a runner that continues produce IDENTICAL
# output, so the witness would pass on a runner doing the opposite of
# the stated policy. The sentinel after the failure, asserted to have
# written its own log, is the only thing that separates them.
#
# `true` and `false` are the entire workload, so this stays on the
# cheap side of the suite. Whether `cargo build` really fails is
# cargo's business; whether THIS SCRIPT names the right gate when a
# command fails is the criterion, and that is orthogonal to which
# command failed.
# ---------------------------------------------------------------------
emit_self_test_plan() {
printf 'self-pass\ttrue\n'
printf 'build-crdt\tfalse\n'
printf 'self-sentinel\ttrue\n'
}
# ---------------------------------------------------------------------
# Pruning.
#
@ -329,6 +410,7 @@ while [ $# -gt 0 ]; do
--print-target-dir) MODE=printdir; shift ;;
--init) MODE=init; shift ;;
--prune) MODE=prune; shift ;;
--self-test) MODE=selftest; shift ;;
--force) FORCE=1; shift ;;
-h|--help) usage ;;
*) echo "gate: unknown argument: $1" >&2; usage ;;
@ -397,11 +479,23 @@ echo "gate: logs $LOGDIR"
# trap removes it, so it should be gone once the run finishes.
echo "gate: ambient $AMBIENT"
[ -n "$ACCEPTANCE" ] && echo "gate: acceptance $ACCEPTANCE"
[ "$PROTOCOL" = 1 ] && echo "gate: protocol yes (CRDT workspace sweep added)"
[ "$PROTOCOL" = 1 ] && echo "gate: protocol yes (CRDT build + workspace sweep added)"
if [ "$MODE" = selftest ]; then
echo "gate: SELF-TEST hardcoded synthetic plan --- NO real gate runs."
echo "gate: the middle step fails ON PURPOSE, so a non-zero"
echo "gate: exit is this mode working, not this mode broken."
fi
echo
# The self-test hands the REAL runner loop below a different plan file.
# Everything after this point is shared, which is the point: a witness
# that exercised its own copy of the runner would witness nothing.
PLAN_FILE="$LOGDIR/plan.txt"
emit_plan > "$PLAN_FILE"
if [ "$MODE" = selftest ]; then
emit_self_test_plan > "$PLAN_FILE"
else
emit_plan > "$PLAN_FILE"
fi
N=0
FAILED=''

View File

@ -127,6 +127,156 @@ fn the_crdt_workspace_sweep_is_added_by_protocol_and_absent_without_it() {
);
}
/// **The precondition the plan did not encode**, and the reason a green
/// `--protocol` run could mean nothing.
///
/// The crdt workspace sweep spawns `pmacs-gpu` as a *process*, and no
/// `cargo test` run produces that binary — `pmacs-gpu` has no `tests/`
/// directory, so cargo never uplifts its bin to `debug/pmacs-gpu`. On a
/// cold target directory the sweep fails twelve
/// `gpu_invocation_acceptance::crdt::*` tests on *"build pmacs-gpu
/// before this acceptance suite"*. Before per-worktree target
/// directories (#225) every worktree shared one that nearly always
/// already held the binary, so the precondition was satisfied **by
/// accident** — and the hazard was never the red gate, it was a green
/// one decided by the build directory rather than by the diff.
///
/// **The exact command is asserted, not just the step's name and
/// position.** A `build-crdt` running plain `cargo build` would sit in
/// the right place under the right name and leave the gate exactly as
/// unsound: the crdt sweep needs *those* features, and the wrong ones
/// produce a binary the sweep cannot use.
#[test]
fn the_crdt_sweep_is_immediately_preceded_by_the_build_that_produces_its_binary() {
let root = tempfile::tempdir().expect("tempdir");
let build = "cargo build --workspace --no-default-features --features luajit,crdt";
let crdt_sweep = "cargo test --workspace --features crdt --no-fail-fast -- --skip basedpyright";
let (plan, err, ok) = run(root.path(), &["--protocol", "--print-plan"]);
assert!(ok, "--protocol --print-plan must succeed; stderr:\n{err}");
let b = plan
.find(build)
.unwrap_or_else(|| panic!("the crdt sweep's build is missing; plan was:\n{plan}"));
let s = plan
.find(crdt_sweep)
.unwrap_or_else(|| panic!("the crdt sweep is missing; plan was:\n{plan}"));
// IMMEDIATELY before: one newline between them and nothing else. A
// build that merely appears *somewhere* earlier could be separated
// from the sweep by a step that rewrites the same target directory.
assert_eq!(
&plan[b + build.len()..s],
"\n",
"the build must run IMMEDIATELY before the crdt sweep; plan was:\n{plan}"
);
}
/// **Conditionality, settled by measurement rather than by reading** —
/// which is the whole methodological point of this lane, since the
/// defect it repairs was a precondition nobody checked.
///
/// Measured 2026-08-09 on a disposable target directory, with
/// `debug/pmacs-gpu` asserted **absent** before each run and each sweep
/// run alone from that same cold state: the default sweep exited **0**
/// and left `debug/pmacs-gpu` **still absent** — it never builds the
/// binary and never needs it — while the crdt sweep exited **101** with
/// exactly twelve `gpu_invocation_acceptance::crdt::*` failures.
///
/// So an unconditional build would be a real cost paid for nothing on
/// every ordinary lane.
#[test]
fn the_crdt_build_is_absent_without_protocol() {
let root = tempfile::tempdir().expect("tempdir");
let (plan, _, ok) = run(root.path(), &["--print-plan"]);
assert!(ok, "--print-plan must succeed");
assert!(
!plan.contains("cargo build"),
"the default sweep passes on a tree with no pmacs-gpu at all, so a \
normal lane must not pay for a workspace build; plan was:\n{plan}"
);
}
/// **The attribution and continuation criteria, made observable.**
///
/// Everything else in this file drives a no-gates path, so it can prove
/// a step's name and its order and **nothing** about what the runner
/// does when a step fails. `--self-test` closes that gap by handing the
/// *real* runner loop a hardcoded three-line plan — a passing step, a
/// failing one named `build-crdt`, and a passing sentinel after it.
///
/// **Why `build-crdt` must be its own step** is exactly what this
/// witnesses: folded into the sweep as `cargo build … && cargo test …`,
/// a *build* failure would be reported under the name `sweep-crdt` — a
/// wrong attribution in the one place this script exists to be
/// trustworthy about.
///
/// **The sentinel assertion is the load-bearing one.** With the failure
/// last, a runner that aborts and one that continues produce identical
/// output, so a two-line witness would pass on a runner doing the
/// opposite of the stated `--no-fail-fast` policy. The sentinel's own
/// log existing is the only thing that separates them — delete that
/// assertion and this test stops testing continuation at all.
///
/// The plan is a literal inside the script on purpose. Making
/// `PLAN_FILE` injectable would let this test supply its own commands,
/// and would turn the runner's `eval` into a general command executor —
/// the same defect the `--acceptance` refusal above exists to prevent.
#[test]
fn self_test_names_the_failing_gate_and_the_suite_continues_past_it() {
let root = tempfile::tempdir().expect("tempdir");
let (out, err, ok) = run(root.path(), &["--self-test"]);
assert!(
!ok,
"a plan containing a failing step must exit non-zero; stdout:\n{out}stderr:\n{err}"
);
assert!(
out.contains("build-crdt"),
"the failing gate must be named as it runs; stdout:\n{out}"
);
assert!(
err.contains("FAILED: build-crdt"),
"the failing gate must be listed under FAILED: by its OWN name; stderr:\n{err}"
);
// The runner claims a log path for the failure. Assert the file is
// actually there: a tool that prints a path it did not write is
// worse than one that prints nothing, because the absence is only
// discovered while chasing a real failure.
let claimed = err
.lines()
.find_map(|l| l.split_once("log: ").map(|(_, path)| path.trim()))
.unwrap_or_else(|| panic!("the failing gate's log path must be printed; stderr:\n{err}"));
assert!(
claimed.ends_with("02-build-crdt.log"),
"the log must be numbered and named for the gate that failed; was {claimed}"
);
assert!(
Path::new(claimed).is_file(),
"the runner must WRITE the log it claims at {claimed}"
);
let logdir = Path::new(claimed)
.parent()
.expect("the log lives in a log directory");
assert!(
logdir.join("01-self-pass.log").is_file(),
"the step before the failure must have its own log; dir was {}",
logdir.display()
);
// THE ASSERTION THE WHOLE MODE EXISTS FOR.
assert!(
logdir.join("03-self-sentinel.log").is_file(),
"the suite must CONTINUE past a failed gate — the sentinel after \
build-crdt wrote no log, so this runner ABORTED. Stdout:\n{out}"
);
assert!(
out.contains("self-sentinel"),
"the sentinel must be reported like any other gate; stdout:\n{out}"
);
}
/// The seam handoff §3 keeps authority over: a script cannot infer
/// which acceptance suites a change touched, so it runs what it is
/// handed — each one, in order.