Four findings on revision 3, all upheld. Answering finding 1 turned up
something that reframes the lane.
THE ONSET. sweep-crdt appears SEVENTEEN times in this target directory's
gate logs. The ctrl_c failure appears in exactly the LAST THREE, and the
test passed --- both copies, "... ok" --- inside the stage before them.
Last green 20260815T185708Z, first red 20260816T063330Z, no reboot
between. The three earlier red sweeps failed on unrelated rows. So
"pre-existing on main" holds (F1 at 72da24a reproduces it) but "always
broken" was never established and is now contradicted. D0 gains a first
part: bisect that window. A test that passed fourteen times in this
stage and then failed three times running has a change behind it, and
that is worth more than further reduction --- which has isolated
nothing.
1. Both ledgers still carried the falsified R9 conclusions. This branch
listed --workspace unification and preceding tests as ruled out
while the section above described an interaction; said "five call
sites" immediately before correcting to six; and labelled the
framing revision 2. The held branch was worse: --workspace refuted,
R9 "same binaries", later packages not implicable, cause cumulative
across 37 binaries. All corrected and pushed (5b9abd8). §11 no
longer asserts the held lane is clean; it records a re-verified
checklist, since asserting that prematurely is what went wrong.
2. Manifest now carries complete argv for R7-R9 and F5 --- abbreviations
are not reconstructable invocations. F5 is disambiguated: the
framing cited gate ...-2144707 while the manifest cited ...-2375685,
two distinct real runs. Enumerating them gives F1-F7: the red count
is SEVEN, not five, each with its own log digest. F5 also carries an
extra failing binary the others do not.
3. "Workspace artifact family" conflated Cargo suffix with byte
identity and is withdrawn as a grouping. Demonstrated: F1 in the
main worktree executed the same suffixes -5d9105cb and -d4dae4f0,
but the bytes there are e0578039/00f06aeb versus the panel
worktree's 1b3cc86c/ede0c07d. Each run now records the suffix its
log shows and byte identity as UNKNOWN, since target dirs have been
overwritten and a hash computed today is not the hash that ran.
4. The interaction table is demoted to a description of what was
observed. Revision 3 disclaimed its inputs and then asserted a
finding from them, which cannot both hold. A3 no longer speaks of an
established "R9 paradox" --- there is none to explain, because the
comparison was never made; it requires D0 to recreate it first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Five findings on revision 2, all upheld. The first invalidates its
strongest claim.
1. R9 executed gpu_initial_target_acceptance-91f51d0b and
gpu_invocation_acceptance-6b4b8223; the failing sweeps executed
-5d9105cb and -d4dae4f0. Verified byte-different by sha256. Cargo's
target selection changes the fingerprint, so command shape changes
the executable. "Same binaries" is now "same target names and
order". What the evidence supports is an INTERACTION --- prior
targets alone green (R9), workspace artifacts alone green (R10),
both together red (F1-F5) --- so --workspace selection is not
sufficient by itself and NOT ruled out. The claim that other
packages "cannot be implicated" because their targets run after the
failure is withdrawn: later-selected packages can affect the build
graph and fingerprints before their tests ever run.
2. Both ledgers made internally consistent and portable. This branch's
asserted default-disposition death and then withdrew it further
down; the assertion is gone. panel-mapping-generation still carried
"119 binaries green one red", the >=8s arithmetic, the default-action
claim and the >6s selector --- corrected on its own branch and pushed
at 779a6bd.
3. Provenance is now a pushed document, docs/probe-sigint-evidence.md:
exact command, worktree, HEAD, cleanliness, artifact family, result
and log digest per physical run. R1 and R2 have no preserved log,
and revision 2 double-counted one log as both R2 and R6. Cleanliness
is UNKNOWN for every pre-manifest run and is not inferred. R1-R10
ran in the panel-mapping-generation worktree, not at main. D0 now
precedes every other diagnostic: re-run the matrix at main under a
harness capturing provenance AND the artifact hashes executed.
4. "The probe never blocks indefinitely" narrowed to "the event loop
wakes at least every 50ms". The stdin reader blocks in read_to_end
(:1109) and, once ready, the loop leaves only when stdin closes
(:1212), so the process is not bounded.
5. Launcher call sites: six under --features crdt (:509 :534 :544 :574
:725 :1097, inside #[cfg(feature = "crdt")] mod crdt). The other two
--gpu arguments are under #[cfg(not(...))] and compiled out.
Revision 2 said five while citing eight.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Revision 1 rejected on five findings, all upheld.
1. The >6s selector could not have captured the failure. Both
reproducing binaries finish in ~5.19s INCLUDING the 5s timeout
(:3097, :3131), so the failing launcher lives about 5.1s. This also
falsifies my earlier retraction, which had argued the instance "must
live >=8s" --- so the "mechanism located" claim is NOT refuted by
that argument. It stays unproven for a different reason: the suite
spawns launchers from five call sites, so command line alone cannot
attribute one to this test. Key on the PID the test records.
2. Diagnostics rewritten to DISCRIMINATE blocked delivery, inherited
ignore, and an escaped process group: before-and-after snapshots for
test parent / launcher / probe, per-thread SigBlk from
/proc/<pid>/task/*/status, SigPnd/ShdPnd, and PID/PPID/PGID/SID.
Relatedly, "two processes with default disposition" is withdrawn ---
SIG_IGN is inherited across fork and survives exec, so absence of
handler code says nothing about runtime disposition, and inherited
ignore is the leading hypothesis precisely because the source is
silent. Revision 1 contradicted its own hypothesis.
3. Counts corrected: 119 green result summaries and TWO red binaries,
not "119 binaries green, one red". Reductions are now enumerated
R1-R10 and F1-F5 with command, run count and log each, preserved off
the tmpfs --- /tmp is a tmpfs and these were nearly lost mid-lane.
4. Acceptance contract corrected: A2 now requires three consecutive
green runs on the reviewed fixed head of this branch, not on main,
which is unobtainable before approval and merge; journey step 12(a)
"closing is clean" is named, since revision 1 reasoned from grade
movement which §20 warns against; and A5 is explicitly conditional
on D4, with bet 1 restated as a bet --- the witness uses a wrapper
and headless probe, not the real GUI path.
5. Portability closed: this branch now tracks
githubsucks/gpu-probe-sigint-teardown, and panel-mapping-generation
was pushed to 16cf3a2 so its retraction travels.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
`ctrl_c_on_launcher_group_does_not_reach_spawned_daemon` fails in gate
stage `sweep-crdt` with "child did not exit within 5s". It is
PRE-EXISTING on main --- 72da24a fails it in a clean worktree with its
own target dir --- so while it reds, no branch can present a green
sixteen-stage gate, main included. §5b is held behind this lane.
Framing revision 1, and it proposes NO FIX, because the mechanism is not
known. What it does instead is fix the shape of the problem so the next
attempt is not another guess:
- Ground truth, cited: neither binary handles signals. `run_gpu`
(src/main.rs:324) blocks in `command.status()` with no handler, and
grepping all of pmacs-gpu/src for signal machinery returns nothing.
The probe polls at 50ms. Two processes with default SIGINT
disposition should both die at once --- this deepens the puzzle
rather than explaining it, and the framing says so.
- Ruled out by measurement, with the method for each: load, tmpfs
(tested by experiment, not argument), leaked daemons, inotify,
--workspace feature unification, and any specific preceding test.
- The reduction paradox stated as the problem's real shape: 5/5 in
the full sweep, 0/N in every reduction, including all 37 preceding
targets plus the suite.
- One retracted claim kept as a warning, because it was mine: the
"mechanism located" report described a healthy teardown. The
sampler behind it caught 394 launchers with a 5s maximum lifetime
while the failing instance must live 8s or more.
The first step is diagnostic only: an instrument keyed on the FAILING
instance --- launchers outliving ~6s --- capturing /proc/<pid>/status
signal masks, since SigIgn survives fork and exec while handlers do not.
Acceptance criteria are written now so the fix cannot quietly become
"make the test pass": a demonstrated mechanism with a mutation-tested
witness, sweep-crdt green three consecutive times, the reduction paradox
explained or recorded as unexplained, and no deadline raised or test
skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai