Commit Graph

500 Commits

Author SHA1 Message Date
Levi Neuwirth 758b985c35
feat(protocol): the mapped panel family --- v25 wire shapes and their pins
SS5b's first implementation commit: the two appended variants, the
version constants, and the pins that hold them in place. No gating, no
key, no replay --- those are the next commits, and the variants are
REFUSED everywhere until their gate lands.

**APPENDED AT THE TRUE END, confirmed by the discriminants.**
`PanelPointer` is 15, `TextInput` 16, `PanelPointerMapped` **17**;
`Present` 0, `Absent` 1, `PresentMapped` **2**. "Beside `Present`" would
have been adjacent insertion, which shifts every discriminant below and
silently re-interprets an older peer's bytes. `mapping_generation` is a
`u64`, last within each variant, documented invalid at zero --- the
value a default-constructed sender produces, so accepting it would let
a peer opt out of the check by sending nothing.

**THE COMPILER NAMED EVERY SEAM.** Four non-exhaustive matches:
`semantic_render`'s declaration accessor now sees through both
families, and the three routing sites REFUSE the mapped variant rather
than unwrapping it to legacy meaning. Refusal is the correct default at
an intermediate commit, not a placeholder --- until the frontend can
prove it negotiated v25 it IS a `<= v24` peer for gating purposes, and
painting first would ship a window in which the band is hit-tested with
no mapping identity at all.

**Five mutations, each biting its own rows:**

  insert `PanelPointerMapped` before `TextInput`
      -> the TextInput pin and the mapped pin. `PanelPointer`'s v23 pin
         correctly SURVIVES: its discriminant did not move, which is the
         "only the pin whose discriminant moved fails" behaviour G0a
         specifies
  insert `PresentMapped` before `Absent`
      -> the Absent pin and the mapped-frame pin
  swap `geometry_epoch` / `panel_epoch`
      -> the exact-bytes assertion, while the round-trip stays green.
         That is the blind spot G0b exists for, and it is why every
         adjacent same-typed field carries a distinct value
  bump the wire version without extending the supported set
      -> both new tripwires and 1a's v6 ladder
  move `ADVERTISED_PROTOCOL_VERSION` to 25
      -> the baseline pin

**Version fallout, enumerated rather than discovered one gate at a
time.** Four acceptance-suite tripwires (`bottom_panel_stage2b_gpu`,
`discovery_stage2` x2, `vterm_stage3`, `statusline_segments`) each say
"a wire bump must be a conscious edit here" and each worked. Rather
than fix them one run at a time I grepped the tree for version
assertions and updated all four in one pass.

Review folded five further corrections, two of which fix reasoning of
mine that was wrong:

  - I claimed reversing `frame` and `mapping_generation` "fails to
    compile" because they are different types. **False for NAMED
    variant fields** --- the initializer uses names, so reordering the
    declarations compiles and shifts postcard's positional bytes
    silently. The pin is the only thing catching that.
  - Ladder loops now track `PROTOCOL_VERSION` while TRIPWIRES stay
    literal. I had flattened both to `25`. A tripwire is literal so a
    bump is a conscious edit; a ladder must move, or the next bump
    silently stops testing the top rung. G14b is unaffected ---
    `PANEL_MAPPING_MIN_VERSION` stays literal, because there the
    arithmetic is exactly the hazard.
  - `assert!(24 < MIN)` was a compile-time tautology holding for every
    value above 24. Replaced with the literal equality plus
    `assert_ne!` against `TEXT_INPUT_MIN_VERSION`: the mapped family
    must not share v24's gate, or it is admitted on sessions that
    negotiated only `TextInput`.
  - Statusline support loop reaches `PROTOCOL_VERSION`; public protocol
    history records v25.

**CI-red observations are in the LANE LEDGER, not the registry**, and
that is deliberate: `ci-red-signatures.md` here ends at U9 while the
unmerged replay branch already added a U10, so a row from this branch
would duplicate an id or invent one blind --- which this file's own
history records going wrong, two branches' entries merging "without a
conflict, producing duplicate ids across four sites". R7 twice and the
composition budget once, fragments verified, owed to the registry by
whichever branch merges second.

Gates: all eleven green under `env -u TMPDIR` with `--protocol`,
log 20260815T103555Z. Four runs were needed; three were lost to those
two signatures, not to this diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 17:55:21 +02:00
Levi Neuwirth 775a046ce4
docs(framing): fold the mapping-generation closure audit into revision 16
Close the producer, receiver, protocol-family, and gesture-lifecycle
cross-products in one pass. Separate route witnesses from replay effects,
freeze both old boundaries and new variant fields, and record the bounded
review method for the remaining chain.
2026-08-20 17:55:21 +02:00
Levi Neuwirth 137c7fc736
docs(framing): SS5b revision 16 --- the producer rule was self-defeating, and R7 recurs on a docs-only diff
Answers review of 15. Framing only. Three items reverse a rule 15
introduced, and one retracts a mutation that was not a defect.

**THE PRODUCER RULE CONTRADICTED PROACTIVE CANCELLATION.** The daemon
cancels BEFORE emitting the replacement frame, so the frontend needs
only to clear its local latch when that frame arrives, and send
nothing. Revision 15 asked it to emit a cancellation tail or retain the
latch: the tail is redundant --- the daemon would receive a release for
a gesture it has already settled, which is the duplicate release the
latch exists to prevent --- and RETAINING IS ACTIVELY HARMFUL, because
it manufactures a `Drag` under the NEW generation with no accepted
`Down`. That is the exact orphan the section exists to prevent,
produced by the rule meant to prevent it.

Ordering is what makes the simple rule safe: cancel, then emit. The
frame's arrival IS the cancellation signal; no second channel is
needed. Witnessed as `Down` -> key advances -> replacement frame ->
motion and physical `Up` produce no new drag and no duplicate release.

**THE LATCH HAD ONE TRIGGER AND NEEDED FIVE.** Cancellation runs on
every loss of gesture authority: generation advance, `Absent`, panel or
buffer identity change, geometry-epoch change EVEN AT AN UNCHANGED CELL
TOTAL, and detach. And an ordinary accepted `Up` must clear the latch,
or a later invalidation finds a gesture it believes live and
synthesises a duplicate release for a button already up --- the replay
lane's D1/D2 orphan race, arriving from the daemon's side.

**G9b's MUTATION WAS A VALID IMPLEMENTATION, NOT A DEFECT.** Keying the
dedupe by `(mapping_generation, coord)` preserves same-generation
suppression and naturally admits the first motion under a new
generation. Requiring it to fail would have forbidden a correct design.
Replaced with two real defects: compare only the cell and never key or
reset by generation (the first post-change motion is eaten), and reset
on every same-generation repaint (pixel-rate traffic returns).

**"PROJECTED CELL IDENTITY" CONTRADICTED THE STYLING CONTROL** in the
same section. The wire `Cell` derives `PartialEq` over `glyph`, STYLE
and `attachment` (`pmacs-protocol/src/cell.rs:153`), so an identity
keyed on cell equality moves on a pure recolour --- while the stable
controls rule style out. Terminal identity is now glyph and row
TOPOLOGY plus the view anchor, excluding face, style and cursor, with a
same-glyph/different-style control: the row that catches an
implementation reaching for `Cell` equality because it is right there.

P2s: zero-generation rows added in BOTH directions as independent legs
(a valid `PresentMapped` with generation zero must be rejected
atomically; a zero-generation `PanelPointerMapped` must be refused);
G7 split into outbound mapped-frame and inbound mapped-pointer legs,
since its old mutation only withheld the frame; G2's grid rows/columns
and fold-map-content/`fold_projection`-policy composites split; and
SS20 now names journey steps 5 and 8 while stating neither grade
changes --- an auditor scanning for grade movement alone would
otherwise conclude this slice touches no journey.

**AND R7 RECURRED, ON A DIFF THAT IS ENTIRELY DOCUMENTATION.** The
first `--protocol` run of this tree failed the `gpu` step on
`managed_retry_survives_transients_and_uses_the_successful_stream`,
with all three required fragments verified from the durable log
(`20260815T072601Z`). Recorded as R7's FIFTH occurrence.

It carries the strongest tree exclusion the row has had: occurrences 1
and 4 argued "unrelated lane", while this branch cannot be related at
all --- no Rust, no wire surface, no `pmacs-gpu` file. The line moved
to `attach.rs:1728` from `:1680`, which the row already treats as
occurrence-specific rather than a fragment. Isolated rerun green, and
the full gate green on the re-run (271/271 in the `gpu` step) --- which
per this file's rerun rule establishes INTERMITTENCE ONLY, though here
there is no tree change to exonerate.

What five occurrences across three flavors and five unrelated lanes now
support is that the failure is NOT LANE-CORRELATED. That is evidence
about where the cause is not. The retirement condition is unchanged.

Gates: all eleven green under `env -u TMPDIR` with `--protocol`,
log 20260815T073556Z.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 17:55:21 +02:00
Levi Neuwirth fa980ef1ba
docs(framing): SS5b revision 15 --- appended means LAST, and cancellation had a race
Answers review of 14. Framing only. Four of the six reverse something
14 asserted, and a green protocol gate would not have caught any of
them.

**"BESIDE `Present`/`PanelPointer`" WAS POSITIONALLY DANGEROUS.**
"Beside" reads as adjacent, and adjacent insertion shifts every
discriminant below it --- the exact hazard the appended-only rule
exists for. Appended means LAST: `PresentMapped` after `Absent`,
`PanelPointerMapped` after `TextInput`, with a diagram so the next
reader cannot re-derive it wrongly. Field order is stated exactly,
`mapping_generation` is a `u64`, ZERO IS INVALID --- it is what a
default-constructed or half-initialised sender produces, so accepting
it would let a peer opt out of the check by sending nothing --- and the
gate reads `PANEL_MAPPING_MIN_VERSION = 25` rather than a literal.

**BLANKET REFUSAL STOPPED THE WHEEL AFTER ONE TICK.** The first
effective document wheel changes `view_top`, which advances the key, so
the next already-queued tick carries the old generation and is refused:
the panel scrolls once and goes dead until the frontend observes the
new frame. Local terminal scrollback has the same shape.

The discriminator is whether the gesture USES its coordinate.
Coordinate-free gestures --- the document wheel, non-reporting terminal
scrollback --- cannot be mis-aimed by a stale mapping and are EXEMPT.
A child-reported wheel is the opposite case: SGR carries row and
column, so a stale one aims an application action at a cell the user
never pointed at, and it keeps the check. Two-tick witnesses added,
because without them a blanket-refusal implementation passes every
single-event row in the matrix.

**CANCELLATION WAS REACTIVE AND LOSES A RACE.** If the replacement
mapped frame reaches the frontend before the physical `Up`, the
producer resets `pointer_held` and SUPPRESSES THE VERY EVENT that would
have cancelled --- so the daemon is never told, the selection stays
armed, and the child keeps holding its button. It is now PROACTIVE,
triggered by the authoritative key advancing while a gesture is
accepted, and the producer must emit a cancellation tail or retain the
latch rather than clearing first.

That needs state 14 assumed and never specified: an ACCEPTED-GESTURE
LATCH recording whether the `Down` was accepted, whether it reached the
child, and the coordinate, button and encoding a release must match.
Two rules fall out and are ruled here --- a stale `Up` with no accepted
`Down` is INERT, and cancellation NEVER reclaims a controller another
frontend has since taken, because a stale gesture must not steal a live
one's terminal.

**THE EXISTING SCREEN GENERATION CANNOT BE THE TERMINAL KEY.**
`Screen::changed()` bumps from 39 call sites including `SetStyle`,
`Bell`, the tab-stop operations, cursor-only motion and `SetTitle`.
None of those change what a coordinate denotes, so keying on it would
cancel a drag every time the child recoloured a character. A dedicated
terminal mapping revision is defined over projected cell identity,
retained-row identity and the per-view scroll anchor --- with those
five events as explicit STABLE CONTROLS, so a reader who later reaches
for the convenient counter fails a test instead of shipping a cancelled
drag.

**G5'S EFFECTS ARE NOT PROVABLE ON THIS BRANCH**, and 14 claimed them.
`gesture_last_content_cell` and the document/terminal replay exist only
on `panel-pointer-replay` (`pmacs-gpu/src/main.rs:2143` there); the
same struct here is at `:2124` with no such field. The obligations are
split in a table. G5a --- that the key advancing RAISES cancellation
--- stays here on purpose: the trigger is this slice's rule, and moving
the whole row out would leave the proactive ruling with no witness in
the slice that introduces it.

P2s: mutation legs split (wrap vs gutter, terminal content vs
scrollback, G5a-c, G8a/b, G9a/b); G7 given a positive-path mutation;
G11 expanded --- exhaustion must publish `Absent`, clear input
authority, cancel any accepted gesture and LATCH, or a stale panel
stays painted and permanently inert; the v26 correction finished at the
gate and old-peer cells (`:573`); and the SS20 impact statement added
--- hardens an existing panel island, no journey grade changes, no
config, no background work.

**And the pin correction is mine to make: it EXISTS**, at
`src/protocol.rs:1975`, in the ROOT crate's test module rather than
under `pmacs-protocol/` or `tests/` --- which is exactly where I
searched. `message.rs:524` was right and the doubt was wrong; the
contrary claim is removed from both records.

Gates: all eleven green under `env -u TMPDIR` with `--protocol`,
log 20260814T180105Z.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 17:55:21 +02:00
Levi Neuwirth ca816a1917
docs(framing): the panel cell-mapping generation (v25) --- SS5b revision 14
Own branch, own slice, protocol-bearing, runs alone. Framing only; no
implementation. Blocks `panel-pointer-replay`, which blocks GUI arc 1b.

Answers review of revision 13. Every item below reverses or completes
something 13 got wrong.

**GATING IS REFUSAL, NOT FALLBACK.** Revision 13 said a bare
`PanelPointer` from a new peer would be "handled under the old
semantics". That is a BYPASS: it leaves the exact hole this slice
exists to close, reachable by omitting a field. A >= v25 session
sending the legacy event is REFUSED before mutation, and a >= v25
frontend REJECTS a legacy `Present` rather than painting a band it
cannot safely hit-test. Only negotiated <= v24 keeps legacy semantics;
`Absent` stays common to both families.

**ONE AUTHORITATIVE PER-FRONTEND KEY**, used by projection AND inbound
validation, advanced after any mapping mutation and BEFORE the next
inbound pointer is handled --- whether or not anything has rendered.
Comparing against the last EMITTED frame recreates the hole, because a
mutation not yet painted has still changed the inverse mapping.

**STALE TAILS TERMINATE; THEY DO NOT VANISH.** A blanket drop breaks
liveness: a refused `Up` leaves an empty document selection armed with
a stale anchor, and leaves a reporting terminal child HOLDING A BUTTON
FOREVER. Cancellation is now a ruled outcome --- producer latch reset,
daemon selection and click-chain cleanup, and the child's release
delivered at the last coordinate known good. A cancelled gesture is
explicitly not a replayed one: the release is for liveness, and no
selection or scroll effect is applied from the stale event. Stale
BEGINNINGS may still simply drop.

**THE DOMAIN WAS INCOMPLETE.** `view_left` is added, because 1b makes
horizontal scrolling real. "Cursor movement is stable" is now
CONDITIONAL: a cursor move that triggers vertical or horizontal follow
changes `view_top` or `view_left` and therefore does change the
mapping. Terminal panels are ruled explicitly --- their coordinates are
decided by the SCREEN, so output and scrollback movement change the
generation while their buffer revision does not.

**SS5b HAD NO ACCEPTANCE MATRIX.** G1-G11 now cover the foreign edit
before render, every changing and stable domain entry ROW BY ROW, a
selection repaint that must preserve the generation and let a drag
continue, mid-gesture cancellation, v24 and v25 positive controls with
both wrong-family refusals, identical cells across a generation change
still emitting, atomic retention of frame and generation on an invalid
frame, and fail-closed exhaustion. The per-entry enumeration is
deliberate: one aggregate row cannot show WHICH input moved the key,
and a key ignoring `view_left` passes every vertical-only row.

**MAPPED MOTION KEEPS ITS COALESCING TAGS.** A new variant falling
through to the lossless default would put pixel-rate `Move`/`Drag` on a
bounded queue.

**PINS ACCUMULATE.** Revision 13 said the pin "moves", which would
delete coverage of the shape it protects. `PanelPointer` is retained;
exact `TextInput` bytes are added as the previous-final
`FrontendEvent`; the complete nested `PanelFrame(Absent)` bytes are
added as the previous-final `PanelFramePayload`. Recorded honestly: I
could find NO exact-bytes pin for `PanelPointer` anywhere in the tree,
though `pmacs-protocol/src/message.rs:524` says one is in the tests.
Either my search missed it or the doc overclaims; this slice resolves
it either way, since it must add exact pins regardless.

**AND THIS SLICE OWNS THE VERSION CORRECTION.**
`docs/gui-stage1-input-framing.md` now says 1e's `OpenTarget` is
**v26**, with the reason stated at the top. An expected rebase conflict
on `gui-stage1b-pointer-scroll` is not grounds for leaving the
canonical document false --- which is what I argued last round, and it
was wrong. `ADVERTISED_PROTOCOL_VERSION` stays pinned at 20.

Gates: all ELEVEN green under `env -u TMPDIR`, with `--protocol`
(`build-crdt`, `sweep-crdt`), log 20260814T162843Z.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 17:55:21 +02:00
Levi Neuwirth 24e4039eb6
docs(lane): record the one rd_precondition failure, without a cause
The row rd_precondition_validates_the_whole_conformance_set failed once
in a sweep-crdt run on 2026-08-20 and passed on the two sweeps after it.
The message was not captured, so nothing here explains it --- the
occurrence is recorded and the diagnosis is not.

Also withdraws a mechanism I offered for it. I described the test as
spawning 46 concurrent stubs under load; it runs 45 stubs SEQUENTIALLY
plus one intentional nonexistent-path spawn probe, so there is no
concurrency to be pressured and 46 was a miscount. Thirty consecutive
user-run repetitions at load ~10.5 --- 1,350 stub executions --- did not
reproduce it.

Records the standing instruction that a recurrence must capture the
exact case and error before anyone theorises again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 17:31:13 +02:00
Levi Neuwirth f13506caf5
docs: absorb the SIGINT-guard lane --- #241 merged at f8033bc
Records what the lane closed, measured rather than argued.

A6a is closed by measurement: status 1 with no token classifies as a
boundary error, never `ignored`, green on macOS --- the platform whose
shell exits 1 for an exec failure, which is what produced the original
defect and what a status-only ABI could not distinguish.

A7 stops being "satisfied by disclosure". Both macOS flavours exercised
the helper and gate consumers across the full 45-case shared set. The
R-d consumer stays Linux-only, because its test is crdt-gated while the
macOS jobs build without crdt and Test (crdt) is ubuntu-only --- recorded
as an open gap rather than quietly closed.

Also records that the two m4_24_* rows failing locally under crdt do not
reproduce in CI, at this branch or at 72da24a: local-environment
-specific, not a code defect and not this lane's.

panel-mapping-generation is unblocked, and its sixteen-stage gate must
run in the foreground --- the condition its stage 15 always needed.

Per the standing convention this absorption does not advance any
canonical base to its own commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 15:43:46 +02:00
Levi Neuwirth b492426c69
test(sigint): assert the 45 inputs are DISTINCT, and stop claiming every stub carries the sentinel
Two closure gaps.

1. Both suites asserted only `cases.len() == 45`, so the exact
   45-entries-over-43-distinct-inputs regression could recur unnoticed
   --- the one where X3 collapsed into 1/E/empty and X4 into
   0/V/safe/bare, leaving two framing-specified cases silently
   unexercised. shared_cases() now asserts uniqueness over
   (status, stdout, stderr), inside the generator so no consumer can
   forget it. Verified by reverting both payloads to the sentinel: it
   fails naming X3.

2. Comments and ledger still said every stub emits the sentinel, which
   the explicit X3/X4 payloads had made false. They now say the
   BRANCH-DISCRIMINATING cases carry it while X3 and X4 deliberately
   carry their own --- X3 the canonical ignored wording with no token,
   X4 noise --- and that this is what makes them distinct inputs. The
   duplicated `self::`/`super::` explanation left over from the nesting
   fix is reduced to the correct one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 14:22:12 +02:00
Levi Neuwirth fb8a904923
test(sigint): distinct X3/X4 vectors, nesting-safe paths, and the gate I should have run
Five findings. The first was red CI that my local gate could not have
caught.

1. `crate::common` cannot resolve when gpu_invocation_acceptance.rs is
   compiled as a nested module of gpu_initial_target_acceptance.rs,
   where `crate::` is the outer test crate. Now `super::common`, which
   resolves in both modes --- verified by compiling each target
   explicitly. Clippy's `(Some(1 | 2), true)` folding applied too.

   The reason this shipped: plain `./scripts/gate` omits sweep-crdt,
   the only stage that compiles the nested target under crdt, while
   04-lib-crdt builds the lib alone. This lane gates with `--protocol`,
   and the ledger now says so.

2. X3 and X4 had stopped being the cases the framing specifies:
   stub_script() gave every case the same sentinel stderr, so X3 lacked
   the canonical ignored text and X4 was byte-identical to
   0/V/safe/bare --- 45 entries, 43 distinct inputs. Case now carries an
   explicit stderr payload; X3 emits the canonical wording with no
   token, and both consumers assert they never repeat it.

3. The capture-creation-failure row asserted exit, wording and stage
   output but not residue. It now inspects the temporary root before
   its RAII drop and requires it empty.

4. The exact-token test covered safe and error but not ignored, despite
   the ledger claiming all three. The ignored arm now asserts its exact
   stdout, driven through a SIGINT-ignoring shell.

5. The ledger's claim that the status-2 mutation is caught only by the
   dedicated row is superseded --- the sentinel matrix catches it --- and
   the self-referential "this commit" is replaced by bc7d776.

Also records two PRE-EXISTING crdt-only failures found while gating
properly (m4_24_bare_string_glob_stays_relative and
m4_24_d3_fallback_base_is_the_smallest_attachment_dir): they reproduce
in isolation and fail identically at 72da24a, so they are not this
lane's, and no cause is claimed for them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 11:50:34 +02:00
Levi Neuwirth 9321975197
docs(lane): head-exact gate evidence, and the run that was not
Full gate GREEN on the committed head 8802d6a, all 8 stages, log
20260820T072102Z-3009434.

Two provenance corrections recorded rather than smoothed over:

  - The first attempt on that same head failed 07-sweep on
    composition_overhead_under_ten_percent, a perf budget unrelated to
    this lane's surface, green in isolation and already recorded as a
    recurring signature on the panel-mapping-generation ledger. Both
    runs are kept. No cause is claimed for the first --- only that the
    second is the head-exact evidence.
  - The earlier 20260819T190930Z-2647615 run finished about thirty
    seconds BEFORE bc7d776 was committed, so it described the
    implementation tree, not a committed head. It is relabelled
    accordingly rather than left standing as gate evidence for a commit
    that did not yet exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 09:25:48 +02:00
Levi Neuwirth 8802d6a1a2
test(sigint): share the conformance vectors and assert the exact branch
Four acceptance gaps, all upheld.

1. Neither suite distinguished a validated refusal from a boundary
   error. Both exit 2 (and both produce Err in Rust), so comparing exit
   codes or is_ok() let a validator that accepts EVERY status-2 pair
   pass the whole matrix --- the precise defect A6c exists to catch.
   Every stub now emits a sentinel on stderr, and an Outcome enum
   (Safe / ValidatedIgnored / ValidatedError / Boundary) is asserted
   branch-exact: a validated verdict must surface the sentinel, a
   boundary failure must withhold it. Verified: mutating the gate to
   accept any status 2 now fails the MATRIX, where before it only
   failed a dedicated row. Each helper arm's exact stdout token is
   asserted as well.

2. The 45-case set was duplicated in both suites and could drift while
   both still reported length 45. It now lives in
   tests/common/sigint_conformance.rs and both validators consume the
   same vectors.

3. A8 was incomplete --- nothing forced capture-directory creation to
   fail. A bounded row points TMPDIR at a missing directory so
   `mktemp -d` fails, asserting boundary error 2, no stage execution and
   no residue; mutating the failure branch to fall through makes it
   fail. Temporary directories are RAII throughout, replacing the
   keep()-plus-manual-cleanup shape.

4. The R-d comment still claimed a shared helper means the consumers
   "can never disagree" and described status-only behaviour. Both were
   withdrawn by revision 13; the comment now points at the shared matrix
   as what actually keeps them in step.

36 gate rows, 16 GPU rows, clippy clean, full gate green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-20 09:15:47 +02:00
Levi Neuwirth bc7d776569
feat(gate,test): implement revision 13 --- the validated (status, token) pair
The helper now emits its verdict token on stdout with diagnostics on
stderr, and both consumers validate the PAIR rather than the status
alone. This closes the macOS defect CI found: a shell that cannot
execute the helper exits 1, which the status-only ABI read as
`ignored`, so a broken guard told the operator their environment
ignores SIGINT.

Gate (shell consumer):
  - guard-local capture directory, created before the gate's own
    temporary roots exist, with cleanup armed BEFORE the helper runs and
    disarmed on the safe path so the gate's later trap is undisturbed;
  - `|| sigint_status=$?` retained --- a bare invocation dies under
    `set -eu` before the status is read, which was the original bug;
  - `expected_token` selected by an explicit status case before any
    `set -u`-sensitive use, since an out-of-range status has none;
  - byte comparison via `cmp` against both permitted encodings, because
    a shell variable neither preserves NUL nor carries the child status;
  - the helper's stderr is surfaced ONLY for validated verdicts; a
    boundary failure prints the gate's own wording and withholds the
    untrusted child output;
  - every refusing branch prints status= and token=.

R-d (Rust consumer) validates the same pair from Command::output()
bytes. It needs no capture files, and its spawn-error path has no status
at all --- the boundary the shell cannot represent.

Conformance: 45 shared cases generated as a cross-product over token
class, encoding and status, run by BOTH validators so they cannot
diverge, plus Rust's X2 for 46 overall. 34 gate rows, 16 GPU rows, full
gate green.

Mutations, each biting its row: accepting any status 2 regardless of
token; surfacing child stderr on a boundary failure; emitting the token
to stderr. The first is caught by the dedicated error row rather than
the conformance set --- most of the set's boundary cases have empty
stderr, so they cannot tell which branch produced the exit 2 --- and
that limitation is recorded rather than left implicit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 21:14:30 +02:00
Levi Neuwirth 2a6625ddb9
docs(framing): record revision 13 approval
Revision 13 is approved at 5dece3e after closing the status-preserving
capture, guard-local cleanup, exact-byte grammar, stderr trust, complete
pair matrix and consumer-specific boundary blockers.

The replacement may now be implemented under the A1-A8 contract. PR
#241 remains unmergeable until that implementation is complete, gated,
and green on macOS.
2026-08-19 20:52:25 +02:00
Levi Neuwirth c3ad66f578
docs(framing): revision 13 round 3 --- untrusted stderr, exact pairs, one grammar
Four blocking issues, all upheld. The first defeats the whole design if
left standing.

1. Boundary errors trusted unvalidated stderr. A helper exiting 1 with
   NO token but the canonical "SIGINT is ignored" text would classify
   as boundary error --- correctly --- and then tell the operator their
   environment ignores SIGINT. A6 satisfied in the classification,
   violated in the message actually read. Now: a validated pair's
   stderr IS the diagnosis and is surfaced unchanged; a boundary
   failure's stderr is untrusted, and the consumer emits its own
   wording, omitting the child's or labelling it untrusted. New A6b
   witnesses exactly that case (conformance row 23), with a mutation
   for a consumer that surfaces it anyway.

2. The matrix did not prove exact-pair validation: no invalid status-2
   pair existed, and the expected column collapsed validated
   (2, :error) with boundary errors, so a validator accepting every
   status 2 passed all twelve rows. The matrix is now a 23-case
   cross-product distinguishing `error (validated)` from
   `error (boundary)`, with (2, missing), (2, :safe), (2, :ignored) and
   (2, unknown-version) all boundary. New A6c pins it.

3. Normalisation was internally inconsistent and not implementable
   identically. "Strip one newline then trim ASCII whitespace" removes
   further newlines, so TOKEN\n\n would have validated while the same
   clause demanded single-line output --- and POSIX $() strips ALL
   trailing newlines while Rust returns raw bytes, so the consumers
   could not have agreed even on a correct rule. Replaced by one byte
   grammar, stdout := TOKEN | TOKEN LF, with NO trimming, plus the
   shell sentinel idiom `out=$(helper; printf x); out=${out%x}` so the
   shell preserves what it must compare. Vectors added for extra
   newline, leading newline, surrounding spaces, CRLF and doubled
   token.

4. The ledger's old A7 assertion --- satisfied by disclosure, Linux-only,
   no non-Linux unix reachable --- contradicted its own macOS record
   twenty lines above. Marked explicitly as revision-12 history with
   the live record named.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 20:32:54 +02:00
Levi Neuwirth 8b8a692528
docs(framing): revision 13 round 2 --- the algorithm still implemented r12
Five blocking inconsistencies, all upheld. The first was the worst: the
document specified a validated pair and then printed an algorithm that
emits no tokens and a consumer flow that proceeds on exit 0 alone ---
accepting 0 with a missing token, the exact defect revision 13 forbids.

  1. The algorithm now emits exactly one token per arm on stdout with
     diagnostics on stderr; the consumer flow is pair-validation with
     explicit normalisation (strip one trailing newline, trim ASCII
     whitespace, require exactly one line); and the outcome table is
     keyed on pairs, with a fourth row for boundary error including
     macOS's status 1 with no token. `safe` is validated like the
     others --- a status arriving without its token did not come from
     this helper.

  2. A6a is SCOPED TO THE GATE. R-d never sees a shell status: the gate
     goes through /bin/sh, which turns an exec failure into an exit
     status, while Rust's Command returns a spawn error with no status
     at all --- conformance row 12, not row 5. And macOS CI does not
     compile R-d's test, which is crdt-gated while the macOS jobs build
     without crdt. R-d on macOS is unexercised, and the framing says so
     rather than implying coverage.

  3. A7 is restated against measurement. It cannot still say no
     non-Linux unix was tried when macOS ran and went red: five of six
     helper/gate rows pass there, one defect is named, R-d is recorded
     Linux-only, and the remaining portability claim is labelled a
     contract argument.

  4. "Both consumers use the same helper so they can never disagree" is
     withdrawn --- true when the status WAS the verdict, false once each
     consumer validates a pair independently in a different language.
     Replaced by a twelve-case conformance matrix both validators must
     agree on, including the macOS case and a normalisation case.

  5. The token-to-stderr mutation is remapped from A2 to A1/A3, with
     the reasoning recorded: with stdout empty every outcome becomes
     boundary error, which still satisfies A2 as written since A2 only
     requires "not the deadline message". A2 stays broad and A6 pins
     which diagnosis appears.

The ledger is aligned: the mechanism is established rather than
hypothesised, the "stderr prints the raw status" claim is corrected ---
the number appears only in the catch-all, and this failure took the
other branch --- and revision 12 is marked superseded.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 20:27:43 +02:00
Levi Neuwirth 70f0bc960e
test(gate): carry the gate's stderr into the boundary assertion
CI on 916007b: 12 green, 2 red, both macOS Test jobs, and exactly one
row --- gate_maps_an_unexecutable_helper_to_error_not_ignored, left
Some(1) right Some(2). The other five SIGINT rows pass on macOS.

This is the A7 portability finding the review pre-declared, and it is a
real one: the gate returned 1, meaning `ignored`, for a helper it could
not execute --- the exact conflation §7c forbids.

The cause is not established. The leading hypothesis is that the ABI's
1 is ambiguous by construction: 1 means "ignored", and 1 is also a
status shells hand back for assorted failures. On Linux an unexecutable
file yields 126 and the catch-all maps it to 2; if macOS /bin/sh
returns 1 instead, the two cases are the same number at the boundary
and no catch-all can separate them. That would call for verdicts
outside the range shells produce, which is a design change needing its
own revision --- not something to patch here.

This commit only makes the failure self-diagnosing: the assertion now
includes the gate's stderr, which prints the raw probe status it saw.
The first failure could not say which status produced it, because the
message discarded stderr.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 20:01:38 +02:00
Levi Neuwirth 916007b391
docs(lane): immutable checkpoint SHAs, and drop the stale "No PR"
Two ledger findings, both mine.

The lane block still said "No PR" while its own header and a new entry
recorded PR #241.

And the self-referential checkpoint wording had gone false, which is the
same trap as naming a branch's own tip: "this entry's own commit adds
the A6 rows" was true when written at 167d830 and false by d64d300, and
"the entry's own commit adds only the gate record" was 7cef9ca. Every
event now carries its IMMUTABLE sha --- implementation 3206433, A6 rows
and bounded negative path 167d830, factual corrections c9cc8dd, gate
record 7cef9ca, PR record d64d300 --- and only the branch tip stays
symbolic, which is the one pointer that has to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 19:46:32 +02:00
Levi Neuwirth d64d3009d8
docs(lane): record PR #241
Opened from gpu-probe-sigint-teardown into main after the quiet 8/8 gate
on c9cc8dd. Not merged; awaiting review rounds.

Docs-only, per the recording exemption that keeps gate evidence from
recursing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 19:42:09 +02:00
Levi Neuwirth 7cef9ca375
docs(lane): record both gate runs on c9cc8dd --- the red one included
Full gate GREEN on the committed head c9cc8dd, all 8 stages, log
20260819T160220Z-2339958, started at load 3.90.

The preceding attempt on the SAME head is kept rather than dropped. It
failed 04-lib-crdt and 07-sweep on four wall-clock rows --- the
composition budget, the summary-flatten scaling row, dired's 200ms
budget and a lean4 progress notification --- none of which touches this
lane's change. Load average was 49.6 and an unrelated
./verify_task_state.sh run was compiling under a separate toolchain at
/usr/local/rustup, having started about three minutes in and
overlapping precisely the two failing stages.

That overlap is recorded as evidence of WHEN, not proof of WHY. This
lane already retracted one confident environmental attribution, so the
red run was treated as "not valid evidence" rather than explained away,
and the green run on the same commit is what settles it. Had any of the
four failed again on a quiet machine it would have been a real finding
on this branch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 18:07:28 +02:00
Levi Neuwirth c9cc8dd969
docs: two wrong facts --- 33 rows, and approval at 1fc0df6
Both mine, both checkable against evidence already in the repo.

The ledger said 35 gate-acceptance rows. The suite has 33. The 35 was
git_status_stage1_acceptance's result line, which sits immediately
below gate_script_acceptance's in the sweep log; I read the wrong one.
The correction names the misread so the next reader can see how a
transcription from a sweep log goes wrong.

The framing header newly attributed revision 12's approval to 7752bcb.
It was 1fc0df6 --- as the ledger says and as 7752bcb's own commit
message says in its first line. Restored.

The full gate is re-run on THIS commit rather than on the tree that
preceded it; the previous run finished twenty seconds before 167d830
was committed, so it described an uncommitted tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 17:22:01 +02:00
Levi Neuwirth 167d830932
test(gate): witness A6 in both consumers; bound the negative path
Three findings, all upheld.

1. A6 was witnessed only for the helper. Both consumers now have real
   -path rows.

   Gate side, driven through a stub worktree --- a temp git repo holding
   a copy of scripts/gate and a controlled helper --- so the gate's own
   code path runs against each verdict without touching the checked-in
   helper: a stub exiting 2 refuses with the ERROR wording and never
   "SIGINT is ignored"; a NON-EXECUTABLE stub maps 126 to boundary
   error 2 with its own wording. That second case is what the original
   guard got wrong twice.

   R-d side: the precondition is split into sigint_diagnosis() ->
   Result, so the message is testable rather than reachable only
   through a panic in a test that cannot run under the condition it
   describes. The new row asserts safe proceeds, ignored says so and
   says "NOT a teardown defect", error says "could not determine" and
   never "ignored", and an unrunnable helper is undecidable at the
   boundary.

2. The refusal row violated this suite's no-recursion constraint: it
   invoked the ordinary gate, so a regression of the exact `if !` bug
   would have launched eight real gate stages inside the gate suite.
   It now uses --self-test, which drives the same runner over a
   hardcoded synthetic plan, so the negative path stays bounded
   whatever the guard does. under_ignored_sigint() also takes the
   program and arguments POSITIONALLY --- `exec "$@"` --- instead of
   interpolating them into script text, which broke for any path
   containing a space or shell metacharacter, and every path here comes
   from a tempdir or CARGO_MANIFEST_DIR.

3. The portable checkpoint is recorded: implementation at 3206433,
   pushed, signed, clean, full default gate green 8/8 foreground. The
   framing header no longer says implementation "may proceed" --- it
   reports IMPLEMENTED. And docs/agent-handoff.md §3 gains the durable
   rule: never start the gate or cargo test from a shell that ignores
   SIGINT, `setsid nohup ... &` is forbidden, SIG_IGN is inherited
   across fork and survives exec, the gate refuses with no override,
   and scripts/check-sigint-deliverable answers the question directly.

35 gate-acceptance rows, 16 gpu_invocation_acceptance rows, full gate
green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 16:43:37 +02:00
Levi Neuwirth 32064336ee
fix(gate): the guard never fired --- two shell bugs, now covered by tests
Four findings, all upheld, and the first was a live bug I shipped.

1. R-b's non-zero handling was unreachable. scripts/gate runs under
   `set -eu`, so the bare helper invocation killed the shell at exit 1
   or 2 and neither `sigint_status=$?` nor the refusal messages ever
   ran; an unexecutable helper would have escaped as raw 126/127 rather
   than boundary error 2. Reproduced before fixing.

   The first repair was ALSO wrong, and worse: `if ! helper; then
   sigint_status=$?; fi` captures the status of the NEGATED condition,
   which is always 0, so the gate printed the ignored diagnosis and
   then ran the entire suite. The working shape is `helper ||
   sigint_status=$?` --- failure handled, so `set -e` does not fire and
   `$?` is the helper's own --- which is the idiom the helper already
   uses internally. Statuses 1 and 2 pass through unchanged; everything
   else, including 126/127, maps to 2 at the boundary and is never
   reported as "SIGINT is ignored".

   The guard also moved to immediately after the worktree resolves,
   before any log directory, ambient root or tmpdir exists, so a
   refused run leaves nothing behind.

2. The behaviour had no durable coverage, which is exactly why 27
   passing gate tests missed both bugs. Four rows added: helper safe,
   helper ignored, helper error (and never ignored), and gate refusal
   before stage 1. Ignored-SIGINT is simulated with `trap "" INT`,
   which is the real mechanism --- SIG_IGN inherited across fork and
   surviving exec --- not a stand-in. Verified to bite: mutating the
   gate back to either shipped bug fails
   gate_refuses_to_start_when_sigint_is_ignored and nothing else.

3. The ledger now records the implementation, both bugs, the four rows
   and their mutation check.

4. A7 is recorded SATISFIED BY DISCLOSURE, which is the fallback
   revision 12 allows when no non-Linux unix is reachable. The earlier
   "stays open" contradicted the approved contract and is withdrawn.
   Tried: Linux x86_64, all three outcomes, all consumers. Not tried:
   every non-Linux unix. Claimed: POSIX shell only, no /proc, no
   sigaction --- labelled a contract argument, not a measurement.

The full default gate passes all eight stages foreground; it caught a
rustfmt violation in the new test code on the first attempt, which is
the guard-and-gate arrangement working as intended.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 16:19:06 +02:00
Levi Neuwirth 7752bcbf7a
docs(framing): record revision 12 approval
Revision 12 is approved at 1fc0df6 after closing the controlled-arm
provenance, total-helper-ABI, and standing-ledger blockers. Record that
R-b plus R-d implementation may proceed under the replacement A1-A7
contract.
2026-08-19 15:26:50 +02:00
Levi Neuwirth 1fc0df6a8e
docs(framing): close revision 12 approval blockers
Make the second controlled-arm record portable without changing what it
claims: identify head 77b623c, transcribe the actual foreground and
background harness invocations, include the exact evidence-recording
harness, label the captured exit as cargo's, and carry both full binary
digests in both arm columns.

Turn the signal probe into an implementable shared ABI. The checked-in
helper owns classification and diagnostics: 0 is safe, 1 is inherited
ignore, and 2 is probe error. Preserve kill failure in the inner shell,
surface the helper's stderr unchanged in both consumers, and witness the
error outcome in both paths. Correct the mutation mapping so removing
the trap bites foreground success rather than the ignored-signal rows.

Synchronize the active-work ledger with the rerun head, total helper
contract, A1-A7 witnesses, and qualified portability claim.
2026-08-19 15:21:35 +02:00
Levi Neuwirth f607e82263
docs(framing): totalise the helper contract; capture arm digests per run
Three findings, all upheld.

1. The arm provenance was malformed and over-claimed. The "fully
   expanded" background command still contained <the fg command above>
   and <log> placeholders; both table rows were one cell short of the
   header, putting log prefixes under "binary hashes" and leaving the
   digest column empty; and the full binary hashes had been read later
   from reused paths, which cannot retroactively prove what each arm
   executed --- the same provenance rule this document states in §7,
   applied against my own record.

   Rather than weaken the claim, the arms were re-run at head 77b623c
   with FULL SHA-256 captured per run, immediately after each run,
   before anything could rebuild them. Both arms: identical
   0890b78c...4124c and ef6ff1c1...c696, dirty=0, fg exit=0 ok=2, bg
   exit=101 failed=2 SigIgn=0x1007. Byte identity is now carried by the
   capture rather than by inference. Commands are written out with no
   placeholders, and the table cells line up.

2. The ledger still transported superseded operative instructions: a
   "remedy not selected" heading, D0b still owed under A3, journey step
   12(a) still assigned, and the old three-consecutive-run A2 contract.
   All four now match revision 12's §8/§9 --- remedy selected, D0b
   satisfied and not owed, journey steps NONE with gate trustworthiness
   named instead, and A1-A7 replacing the three-run contract, which was
   written for a flakiness that is now explained.

3. The helper contract was not total. The raw probe reaches exit 0 both
   when the kill was a no-op AND when the kill itself failed, so a
   broken probe would report "inherited SIG_IGN" and fail the gate for
   the wrong reason. The helper now owns the classification and returns
   one of safe / ignored / error; consumers consume the verdict and
   never re-derive it. `error` is not folded into `ignored` --- it fails
   the gate with a different diagnosis, because "your environment
   ignores SIGINT" and "the guard could not run" are different
   problems. A6 witnesses the distinct error outcome, A7 requires a
   non-Linux unix exercise or an explicit statement of what was tried,
   and A4 gains a mutation for collapsing error into ignored. R-b's
   stale "needs an explicit override" is reconciled with §7c's no
   -override decision.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 14:58:50 +02:00
Levi Neuwirth 77b623c6ea
docs(lane): the ledger edits ab43132 claimed but did not make
Third occurrence of the same process failure, and the one I had already
written the lesson for twice. ab43132's message said the ledger no
longer claims implementation-absent or mechanism-unknown. The ledger
script died on a stale anchor, and because I separated the steps with a
newline instead of chaining them, `git commit` ran regardless. Gating
one step is not enough when the next step is not gated too.

The ledger now records what the framing does: mechanism KNOWN, remedy
SELECTED as R-b + R-d via the portable probe, A3/D0b satisfied by the
controlled explanation so D0b is not owed, revision 12 awaiting
approval, D1/D2 done rather than "the next step", and the diagnostic
instrument named as the only implementation so far.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 14:47:27 +02:00
Levi Neuwirth 57d8dae511
docs(framing): rewrite the contract §4c had only contradicted
Four findings on revision 11, all upheld.

1. The operative contract still said the opposite of §4c. Bet 1 read as
   open; §7 said the mechanism was unknown with D3/D4 pending; §8 kept
   the old criteria and a conditional A5; §9 claimed a journey-12(a)
   product repair; the ledger and the revision-10 paragraph still said
   D1/D2 had not started. Each is now rewritten as executed, withdrawn,
   discharged or superseded --- §9 in particular now records journey
   steps touched: NONE, for the stated reason that no product behaviour
   changes, with gate trustworthiness named as what the lane does
   affect.

2. The causal evidence is now portable and cleanly reproduced. The
   first capture came from d12.log, which finished five minutes BEFORE
   afe3631 committed the diagnostic code and ran in the reused d0a-B
   target --- inadmissible provenance, now marked as the first sighting
   only. Replaced by controlled arms on committed head 38f2af4,
   dirty=0, in this worktree's own target, with BYTE-IDENTICAL binary
   hashes across arms (0890b78cca22ac1e, ef6ff1c15e11062a): foreground
   exit=0 ok=2, background exit=101 failed=2 SigIgn=0x1007. The outer
   invocation is recorded as a first-class column, since it is the
   causal variable and every earlier "exact command" omitted it. The
   historical foreground/background mapping is marked RECONSTRUCTED
   from the transcript, not captured --- no pre-existing row carries an
   outer-invocation field, which is precisely why the matrix stayed
   confounded for nine revisions.

3. D4 was never executed, so bet 1 is WITHDRAWN BY SCOPE rather than
   falsified, and A5 is RETIRED BY SCOPE rather than struck. Nothing
   here shows a real wgpu session behaves correctly; what is shown is
   that no observed evidence of a user-facing defect survives. The lane
   is now gate/test correctness only.

4. The remedy is not selected. §7b evaluates four candidates --- runner
   normalisation, an early gate guard, fixture isolation via pre_exec,
   and a test-local precondition assertion --- with portability as a
   selection criterion, noting /proc is Linux-only while the suite is
   cfg(unix) and sigaction querying is unsafe. Likely R-b + R-d, but
   nothing is chosen or implemented here. Revision 11's leap from
   "pre_exec is unsafe" to "therefore an assertion" did not follow.

Also renames the meaningless african_close() helper (38f2af4).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 14:35:26 +02:00
Levi Neuwirth f058780a5d
docs(framing): record revision 10 approval
Revision 10 is approved at 4fba9f6 after aligning A3 with the D0b
contingency. The demonstrated D1/D2 mechanism may account directly for
the subset/full difference; otherwise D0b remains mandatory before the
lane closes.

Record that diagnostic-only D1/D2 are authorised but have not started.
No mechanism or fix is claimed yet.
2026-08-19 14:02:58 +02:00
Levi Neuwirth 5f5fde6dde
docs(framing): revision 10 --- awaiting approval; fix the corrupted provenance
Two findings, both upheld.

1. The portable provenance was corrupted and incomplete --- worse than
   the machine-local pointer it replaced, because it looked verifiable
   and was not. Every log digest had lost its leading hex character
   (A#1 recorded as 1c0fe47d55d8f5e... where the value is
   e1c0fe47d55d8f5e): the extraction started one byte late in
   `logsha=<value>`. The captured /tmp and MemAvailable columns were
   dropped, and the command block used ellipsed paths. All ten digests
   are corrected, both columns restored, and the command is written out
   in full with only two named placeholders.

   Separately: `uptime` was NEVER CAPTURED. §7's condition list names
   it; the harness kept the load averages from it and discarded the
   elapsed time. It is now recorded as UNKNOWN for all ten runs, with
   the condition list marked as only partially satisfied rather than
   implied met. The classifications stand --- none depends on uptime ---
   and D1/D2's harness must capture the whole list.

2. Retiring D0b materially changes the approved diagnostic sequence,
   which made D0b mandatory before every other diagnostic. The document
   still claimed revision 9, approved at 15c25ec, for a decision that
   approval does not contain. Promoted to revision 10 and marked
   AWAITING APPROVAL; D0a's execution and result are reported under
   revision 9, and D1/D2 do not begin until revision 10 is approved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 13:39:46 +02:00
Levi Neuwirth 18b74d7a97
docs(evidence): narrow the D0a conclusion; retire D0b as a precondition
Three findings, all upheld.

1. The causal conclusion overreached, in the same way this lane has
   overreached before. Uniform-red at both endpoints today proves only
   that the two commits DO NOT DISCRIMINATE UNDER CURRENT CONDITIONS.
   "Source hypothesis eliminated", "the interval cannot contain the
   transition" and "unreachable by source" are withdrawn from the
   framing, the manifest and the ledger: a historical regression could
   be masked by a later environmental effect, or by a source/environment
   interaction under which both commits now fail. Failing to
   discriminate is not the same as not differing. "No bisect is
   justified under current conditions" is what survives, and the
   approved endpoint table's two uniform-same rows are corrected to say
   the same thing.

2. D0b was still mandatory, and going to D1/D2 would have skipped an
   approved step. It is now RETIRED AS A PRECONDITION with the reason
   recorded: it existed to make the reduction matrix trustworthy so the
   subset-vs-full comparison could locate the mechanism indirectly,
   and D0a has since produced a reliable direct reproduction that D1/D2
   measure against. Re-running ten reduction rows to sharpen an
   indirect instrument while a direct one is in hand is the wrong order
   of work. The obligation is NOT discharged: A3 still binds, so if
   D1/D2 fail to account for why every subset passed, D0b runs before
   this lane closes.

3. Provenance is now portable. The exact per-run command and a
   transcribed ten-row table --- start time, class, red bins, load,
   freeMB, daemon count, log digest --- are committed, rather than
   delegated to a machine-local results.tsv. Raw logs stay local by
   design. The transcription also surfaces something the delegation hid:
   the leaked-daemon count climbs 72 -> 108, four per run, monotonically
   while every run classifies identically. Recorded, not implicated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 13:23:19 +02:00
Levi Neuwirth 24a84b5381
docs(evidence): D0a executed --- the source hypothesis is eliminated
Ten runs under the approved contract: counterbalanced A B B A A B B A A
B, N = 5 per endpoint, clean detached worktrees at 7599661 and 724b785,
isolated target directories, the gate's build-crdt precondition then its
sweep-crdt command, dirty=0 verified per run. Zero voids, zero splits.

A (7599661) uniform-red. B (724b785) uniform-red. By the approved
endpoint table that is the both-endpoints-uniform-same row: the
difference is NOT captured by those two commits.

What it settles:

  - No bisect of 7599661..724b785 is justified, and none will run.
    7599661 passed inside sweep-crdt on 08-15 and fails 5/5 clean today,
    so the interval cannot contain the transition.
  - The onset window is demoted --- still a true observation, but not
    reachable by source.
  - A RELIABLE REPRODUCTION now exists: 10/10 today across two commits
    at ~4 minutes per run. This is D0a's most useful product, because
    D1/D2 no longer depend on catching a rare event.

What it does not settle: anything about the mechanism. One cheap
negative on "what else changed" --- no package activity in the window per
pacman.log, nearest on 08-18 --- and it is not pursued further, because
with a reproduction in hand direct measurement dominates archaeology.

A's three extra failing binaries are recorded rather than swept up:
a54_real_daemon_real_pty_and_headless_gpu_render..., a v21/v20 row
expected to differ at that older commit, and m6_1_pty_mode_lifecycle.
Two of the three are process/PTY-spawn rows, the same family as the
target. None affect classification, which reads only the two target
copies.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 13:04:59 +02:00
Levi Neuwirth bdef05cd02
docs(framing): record revision 9 approval
Revision 9 is approved at 15c25ec after the portable manifest and compact
ledger summary preserve the endpoint direction required by D0a.

Record that approval durably before diagnostic implementation begins. The
mechanism remains unknown, no fix is proposed, and panel-mapping-generation
remains held until this teardown lane closes.
2026-08-19 12:14:21 +02:00
Levi Neuwirth 15c25ecaad
docs(evidence): preserve endpoint direction in D0 summary
The portable manifest collapsed the two clean-split directions even though
the governing endpoint table permits a bisect only when 7599661 is uniform
green and 724b785 is uniform red. Preserve that direction explicitly, and
carry the same distinction in the compact active-work summary.

The inverted split remains a real difference, but it contradicts the onset
reading and therefore requires that reading to be re-examined before any
bisect.
2026-08-19 12:02:21 +02:00
Levi Neuwirth 74dbd342a6
docs(framing): revision 9 --- a total classifier, and honest counterbalancing
Two D0a findings on revision 8, both upheld.

1. The classifier was not total. "Clean split" and "mixed" left five
   outcomes unprescribed, and two of them are in the historical logs
   already: 20260815T182846Z-708693 died compiling pmacs so neither
   copy executed, and ...-2839374 / ...-830195 were red on unrelated
   rows while both ctrl_c copies passed.

   A run is now classified from THE TWO COPIES OF THE TARGET TEST and
   nothing else --- green (both ok), red (both FAILED), split (copies
   disagree), void (either did not execute). A sweep red only on
   unrelated tests is therefore a green run, with the unrelated
   failures recorded as evidence about environment stability. A split
   STOPS the procedure, since two copies of one source disagreeing
   within a run is its own defect. Voids are discarded and re-run on a
   budget of 3, after which the environment is too unstable to classify
   anything and D0a stops.

   Endpoint verdicts are uniform green, uniform red, or mixed, and a
   six-row table prescribes every combination: clean split permits the
   bisect; an inverted split is a real difference that falsifies which
   endpoint was believed good; both-uniform-green and both-uniform-red
   each mean the difference is not captured by those commits; mixed at
   either endpoint means intermittency under fixed source and forbids a
   bisect. The manifest had attached "difference is not captured" to
   the mixed case --- that conclusion belongs to the uniform-same rows,
   and is moved.

2. Strict A/B/A/B does not make drift "hit both arms equally": B always
   follows A and owns the final time point. Runs are now counterbalanced
   AB BA AB BA AB, which removes systematic order confounding; the
   residual last-slot asymmetry is accepted and stated rather than
   claimed away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 10:23:43 +02:00
Levi Neuwirth 4e84ff0050
docs(framing): revision 8 --- the superseded one-run rule was still in force
Four findings on revision 7, all upheld.

1. The old one-run D0 rule survived in three durable places --- the
   manifest, this branch's ledger, and the framing's own §4a --- each
   still permitting a bisect when the endpoints merely "differ". That
   contradicts the N = 5 clean-split contract added in revision 7. All
   three now defer to that contract, and §4a's "needs only that the two
   clean endpoints differ now" is marked as the superseded rule it is.

2. D0a still overstated its evidence, in three ways now fixed:
     - "context-sensitive by construction, appearing only in the full
       sweep" is downgraded to what has been OBSERVED so far;
     - the historical 7/7 and 13/13 are stated as NOT endpoint-specific
       rates --- of seven reds only F6 ran at 724b785, of the greens only
       the last at 7599661, both with unknown cleanliness;
     - five runs are named a PREDEFINED EVIDENTIARY THRESHOLD chosen so
       the outcome cannot be argued after the fact, not something that
       mathematically separates intermittency.
   And the bisect now specifies its own classifier: every intermediate
   commit uses the identical N = 5 protocol, and a mixed classification
   ABORTS the bisect rather than being guessed, skipped, or rerun until
   it agrees. A bisect with cheaper steps than its endpoints would
   inherit the weakness the contract exists to remove.

3. The artifacts column is now exact per run, read from each log:
   R1/R2 UNKNOWN (no log preserved), R3 -5d9105cb/-d4dae4f0, R4 and R5
   -6b4b8223 only, R6 -91f51d0b/-6b4b8223. R8's citation was half2.log:1;
   the executable lines are 438 and 459. The framing's last "not same
   binaries" is now "not the same compilations".

4. (Held ledger, 5274d6b.) It named a stale ledger tip and two different
   framing revisions on consecutive lines.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-19 10:08:11 +02:00
Levi Neuwirth 7110256956
docs(framing): revision 7 --- the ancestry supports no causal claim at all
Five findings on revision 6, all upheld.

1. The ancestry pair supports nothing causal. Revision 6 had already
   retreated to "outcome is not determined by commit alone"; that is
   withdrawn too, because different commits CAN deterministically
   produce different outcomes --- this document's own fix-then-regression
   scenario is an example. The two observations differ in commit AND
   environment AND time, so they are simply NON-COMPARABLE. The held
   ledger's "no source-monotonic cause does that" goes with it.

2. D0a was not a valid decision procedure: one unspecified run per
   endpoint cannot establish a regression for a failure that only
   appears in the full sweep. Now specified --- N = 5 full sweep-crdt
   runs per endpoint, INTERLEAVED A/B/A/B so session drift hits both
   arms, identical captured conditions including uptime/free//tmp/
   leaked-daemon count, and a bisect permitted ONLY on a clean split.
   A mixed result means intermittency under fixed source, and no bisect
   is justified at all.

3. "Neither binary contains signal-handling code" is FALSE. The pmacs
   binary does: install_signal_handlers (src/daemon.rs:628) registers
   SIGINT and SIGTERM; it is simply not on run_gpu's path. A grep of
   project sources also cannot exclude a runtime or dependency
   installing a disposition. The established fact is narrow --- no
   explicit installation on run_gpu's path --- and "whatever disposition
   they hold was inherited" is restored to a HYPOTHESIS that D2 must
   measure.

4. Artifact wording finished: no "artifact family", "reduction/
   workspace artifacts" or "different binaries" remain. Every manifest
   row now carries its exact Cargo suffixes read from its log, with a
   stated caveat that those logs are machine-local and this manifest is
   the portable transcription of them.

5. Held ledger pointed at revision 5; it now points at revision 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 18:03:41 +02:00
Levi Neuwirth e2084fbba5
docs(framing): revision 6 --- the ancestry argument shows less than claimed
Three findings on revision 5, all upheld.

1. The ancestry argument overreached. 72da24a failing today while its
   descendant 7599661 passed on 08-15 shows exactly one thing: outcome
   is not determined by commit alone, since the observations come from
   different environments at different times. Revision 5 said a source
   cause was "positively discouraged", that the ancestry "says to
   expect" equal endpoints, and that the change was environmental.
   None follows. It cannot discriminate an environmental change, a
   source/environment interaction, or a fix before 7599661 with a
   regression before 724b785 --- and an ancestor OUTSIDE the interval
   is irrelevant to whether the interval regressed, since a bisect over
   7599661..724b785 needs only that the clean endpoints differ now.

   D0a is unchanged as an action but is now stated as a decision
   procedure with NO predicted outcome: endpoints differ -> bisect that
   interval; endpoints agree -> ask what else changed across the window.

2. The byte-identity withdrawal was incomplete in both ledgers. This
   branch's said the artifacts "are byte-different" and then withdrew
   it two lines later, still said R9 ran "different binaries", and
   still promised an "artifact family". The held ledger still said
   "byte-different" and still called the window a bisect target with
   revision 4's onset conclusion. Both now say "different Cargo
   suffixes/compilations" throughout; historical byte identity is
   UNKNOWN and is never claimed.

3. Provenance slips: R9's observation-table row listed only -6b4b8223
   although it executed both -91f51d0b and -6b4b8223; R10's suffixes
   are at log lines 3 and 24, not 3 and 4; R9's are at 3066 and 3087,
   not 3066 alone. All corrected against the logs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 17:33:23 +02:00
Levi Neuwirth 55053c601a
docs(framing): revision 5 --- the onset is not a source boundary
Four findings on revision 4, all upheld. The third changes what the
lane should do next.

1. Section summaries still carried revision-3 language while the
   manifest carried revision 4's. Framing and ledger now agree: seven
   red runs (F1-F7), not five; the observation table is keyed on
   compilation set rather than an invented "workspace artifact family";
   and it is labelled an observation, not an isolated interaction.

2. The onset count was wrong. Per test copy across the 17 sweep-crdt
   logs: 13 with both copies ok, 1 where NEITHER executed because the
   stage died compiling pmacs (error[E0308]), and 3 with both failed.
   Revision 4's "14 runs, 11 green, 3 red on other tests" mis-stated
   both the count and the kind --- one of those runs never reached the
   test. The two genuinely red-on-other-tests sweeps did execute
   ctrl_c, and it passed.

3. D0a cannot be a source bisect, and the evidence argues against one.
   Reflog and commit times put HEAD at 7599661 during the last green
   (3c06176 landed 40s after it finished) and at 724b785 during the
   first red (5174f73 landed 08:45:41, after that run ended 08:42:01;
   the manifest had recorded F6 at 5174f73, which was wrong).
   Cleanliness was captured at neither endpoint. And 72da24a is an
   ANCESTOR of the passing 7599661 yet fails today --- no
   source-monotonic cause produces that. D0a now reproduces the two
   endpoints CLEAN, in isolated target directories, and a bisect is
   justified only if they differ.

4. Manifest completed: R9 carries full argv rather than a recipe; R7
   lists only gpu_invocation-6b4b8223, since R7 does not select
   gpu_initial_target; R10 lists both -5d9105cb and -d4dae4f0.

Also withdraws "byte-different" everywhere. The bytes a historical run
executed are not knowable --- target dirs have been overwritten, and a
hash computed today is the current occupant's. Three levels are now kept
apart in the manifest: suffix (known), today's bytes at a path (known),
and the bytes a past run executed (UNKNOWN). Differing suffixes mean
differing Cargo metadata hashes, which is enough to void the comparison
and is all that is claimed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 16:06:46 +02:00
Levi Neuwirth 9332d5a616
docs(framing): revision 4 --- and the failure has a datable onset
Four findings on revision 3, all upheld. Answering finding 1 turned up
something that reframes the lane.

THE ONSET. sweep-crdt appears SEVENTEEN times in this target directory's
gate logs. The ctrl_c failure appears in exactly the LAST THREE, and the
test passed --- both copies, "... ok" --- inside the stage before them.
Last green 20260815T185708Z, first red 20260816T063330Z, no reboot
between. The three earlier red sweeps failed on unrelated rows. So
"pre-existing on main" holds (F1 at 72da24a reproduces it) but "always
broken" was never established and is now contradicted. D0 gains a first
part: bisect that window. A test that passed fourteen times in this
stage and then failed three times running has a change behind it, and
that is worth more than further reduction --- which has isolated
nothing.

1. Both ledgers still carried the falsified R9 conclusions. This branch
   listed --workspace unification and preceding tests as ruled out
   while the section above described an interaction; said "five call
   sites" immediately before correcting to six; and labelled the
   framing revision 2. The held branch was worse: --workspace refuted,
   R9 "same binaries", later packages not implicable, cause cumulative
   across 37 binaries. All corrected and pushed (5b9abd8). §11 no
   longer asserts the held lane is clean; it records a re-verified
   checklist, since asserting that prematurely is what went wrong.

2. Manifest now carries complete argv for R7-R9 and F5 --- abbreviations
   are not reconstructable invocations. F5 is disambiguated: the
   framing cited gate ...-2144707 while the manifest cited ...-2375685,
   two distinct real runs. Enumerating them gives F1-F7: the red count
   is SEVEN, not five, each with its own log digest. F5 also carries an
   extra failing binary the others do not.

3. "Workspace artifact family" conflated Cargo suffix with byte
   identity and is withdrawn as a grouping. Demonstrated: F1 in the
   main worktree executed the same suffixes -5d9105cb and -d4dae4f0,
   but the bytes there are e0578039/00f06aeb versus the panel
   worktree's 1b3cc86c/ede0c07d. Each run now records the suffix its
   log shows and byte identity as UNKNOWN, since target dirs have been
   overwritten and a hash computed today is not the hash that ran.

4. The interaction table is demoted to a description of what was
   observed. Revision 3 disclaimed its inputs and then asserted a
   finding from them, which cannot both hold. A3 no longer speaks of an
   established "R9 paradox" --- there is none to explain, because the
   comparison was never made; it requires D0 to recreate it first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 15:50:53 +02:00
Levi Neuwirth 4e1ca68b4c
docs(framing): revision 3 --- R9 did not run the same binaries
Five findings on revision 2, all upheld. The first invalidates its
strongest claim.

1. R9 executed gpu_initial_target_acceptance-91f51d0b and
   gpu_invocation_acceptance-6b4b8223; the failing sweeps executed
   -5d9105cb and -d4dae4f0. Verified byte-different by sha256. Cargo's
   target selection changes the fingerprint, so command shape changes
   the executable. "Same binaries" is now "same target names and
   order". What the evidence supports is an INTERACTION --- prior
   targets alone green (R9), workspace artifacts alone green (R10),
   both together red (F1-F5) --- so --workspace selection is not
   sufficient by itself and NOT ruled out. The claim that other
   packages "cannot be implicated" because their targets run after the
   failure is withdrawn: later-selected packages can affect the build
   graph and fingerprints before their tests ever run.

2. Both ledgers made internally consistent and portable. This branch's
   asserted default-disposition death and then withdrew it further
   down; the assertion is gone. panel-mapping-generation still carried
   "119 binaries green one red", the >=8s arithmetic, the default-action
   claim and the >6s selector --- corrected on its own branch and pushed
   at 779a6bd.

3. Provenance is now a pushed document, docs/probe-sigint-evidence.md:
   exact command, worktree, HEAD, cleanliness, artifact family, result
   and log digest per physical run. R1 and R2 have no preserved log,
   and revision 2 double-counted one log as both R2 and R6. Cleanliness
   is UNKNOWN for every pre-manifest run and is not inferred. R1-R10
   ran in the panel-mapping-generation worktree, not at main. D0 now
   precedes every other diagnostic: re-run the matrix at main under a
   harness capturing provenance AND the artifact hashes executed.

4. "The probe never blocks indefinitely" narrowed to "the event loop
   wakes at least every 50ms". The stdin reader blocks in read_to_end
   (:1109) and, once ready, the loop leaves only when stdin closes
   (:1212), so the process is not bounded.

5. Launcher call sites: six under --features crdt (:509 :534 :544 :574
   :725 :1097, inside #[cfg(feature = "crdt")] mod crdt). The other two
   --gpu arguments are under #[cfg(not(...))] and compiled out.
   Revision 2 said five while citing eight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 15:36:27 +02:00
Levi Neuwirth 9988e974af
docs(framing): revision 2 --- five findings, two of them my own retractions
Revision 1 rejected on five findings, all upheld.

1. The >6s selector could not have captured the failure. Both
   reproducing binaries finish in ~5.19s INCLUDING the 5s timeout
   (:3097, :3131), so the failing launcher lives about 5.1s. This also
   falsifies my earlier retraction, which had argued the instance "must
   live >=8s" --- so the "mechanism located" claim is NOT refuted by
   that argument. It stays unproven for a different reason: the suite
   spawns launchers from five call sites, so command line alone cannot
   attribute one to this test. Key on the PID the test records.

2. Diagnostics rewritten to DISCRIMINATE blocked delivery, inherited
   ignore, and an escaped process group: before-and-after snapshots for
   test parent / launcher / probe, per-thread SigBlk from
   /proc/<pid>/task/*/status, SigPnd/ShdPnd, and PID/PPID/PGID/SID.
   Relatedly, "two processes with default disposition" is withdrawn ---
   SIG_IGN is inherited across fork and survives exec, so absence of
   handler code says nothing about runtime disposition, and inherited
   ignore is the leading hypothesis precisely because the source is
   silent. Revision 1 contradicted its own hypothesis.

3. Counts corrected: 119 green result summaries and TWO red binaries,
   not "119 binaries green, one red". Reductions are now enumerated
   R1-R10 and F1-F5 with command, run count and log each, preserved off
   the tmpfs --- /tmp is a tmpfs and these were nearly lost mid-lane.

4. Acceptance contract corrected: A2 now requires three consecutive
   green runs on the reviewed fixed head of this branch, not on main,
   which is unobtainable before approval and merge; journey step 12(a)
   "closing is clean" is named, since revision 1 reasoned from grade
   movement which §20 warns against; and A5 is explicitly conditional
   on D4, with bet 1 restated as a bet --- the witness uses a wrapper
   and headless probe, not the real GUI path.

5. Portability closed: this branch now tracks
   githubsucks/gpu-probe-sigint-teardown, and panel-mapping-generation
   was pushed to 16cf3a2 so its retraction travels.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 14:47:48 +02:00
Levi Neuwirth f1992a65d9
docs(framing): open the GPU probe SIGINT teardown lane
`ctrl_c_on_launcher_group_does_not_reach_spawned_daemon` fails in gate
stage `sweep-crdt` with "child did not exit within 5s". It is
PRE-EXISTING on main --- 72da24a fails it in a clean worktree with its
own target dir --- so while it reds, no branch can present a green
sixteen-stage gate, main included. §5b is held behind this lane.

Framing revision 1, and it proposes NO FIX, because the mechanism is not
known. What it does instead is fix the shape of the problem so the next
attempt is not another guess:

  - Ground truth, cited: neither binary handles signals. `run_gpu`
    (src/main.rs:324) blocks in `command.status()` with no handler, and
    grepping all of pmacs-gpu/src for signal machinery returns nothing.
    The probe polls at 50ms. Two processes with default SIGINT
    disposition should both die at once --- this deepens the puzzle
    rather than explaining it, and the framing says so.
  - Ruled out by measurement, with the method for each: load, tmpfs
    (tested by experiment, not argument), leaked daemons, inotify,
    --workspace feature unification, and any specific preceding test.
  - The reduction paradox stated as the problem's real shape: 5/5 in
    the full sweep, 0/N in every reduction, including all 37 preceding
    targets plus the suite.
  - One retracted claim kept as a warning, because it was mine: the
    "mechanism located" report described a healthy teardown. The
    sampler behind it caught 394 launchers with a 5s maximum lifetime
    while the failing instance must live 8s or more.

The first step is diagnostic only: an instrument keyed on the FAILING
instance --- launchers outliving ~6s --- capturing /proc/<pid>/status
signal masks, since SigIgn survives fork and exec while handlers do not.

Acceptance criteria are written now so the fix cannot quietly become
"make the test pass": a demonstrated mechanism with a mutation-tested
witness, sweep-crdt green three consecutive times, the reduction paradox
explained or recorded as unexplained, and no deadline raised or test
skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-18 14:22:32 +02:00
Levi Neuwirth da74ae645b
fix(test): the multibyte witness had the wrong axis, and CI proved it
CI went red on three Test jobs with the helper's own panic. That is the
loud failure working as designed --- and the design was wrong.

`${#x}` counting characters is a property of the SHELL first and the
locale second. `bash` counts characters under a UTF-8 locale; `dash`
counts bytes under EVERY locale, having no multibyte handling at all.
`/bin/sh` is `bash` on Arch and `dash` on the Ubuntu runners, so
probing `/bin/sh` and naming a locale established nothing there: no
pair could qualify, and the row could not run.

The helper now probes `(shell, locale)` pairs and the row invokes the
gate THROUGH the qualifying shell. `/bin/sh` is tried first so the real
interpreter is used wherever it qualifies. This is not a contrivance to
get green: `#!/bin/sh` resolves to `bash` on Arch and on macOS, which
is precisely where a `${#VAR}` guard would miscount, and it is the
configuration the guard exists for.

Renumbered, because `M-G-8` was taken. Round 3 assigned it to the
canonical-traversal mutation and the ledger never recorded it, so the
locale exercise reusing the ID was a collision. Canonical `M-G-8` is
restored to the ledger; the locale legs are `M-G-9a-c`. Nine total.

  9a  mutant gate, probed pair -> row fails, boundary row still passes.
      Re-run with /bin/sh EXCLUDED, covering the dash/CI fallback
      path -> still fails.
  9b  SAME mutant gate, pair forced byte-counting -> row passes.
      The defect reproduced rather than argued.
  9c  no pair qualifies -> panic naming shells and locales tried

Record corrections review asked for:

- framing said three rounds and revisions 6a-6c; history is rounds 1-4
  plus this follow-up, and each round is now named for what it fixed
- framing SS2a claimed `${#var}` counts characters under UTF-8 with no
  qualifier --- the same error as the helper's. It now states the shell
  dependence and why the guard measures bytes explicitly.
- the helper's prose said every candidate comes from `locale -a` while
  the code also tried two hardcoded spellings; the doc comment now
  describes what the code does

Gates: all nine green under `env -u TMPDIR`, log 20260813T183646Z.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 20:41:18 +02:00
Levi Neuwirth 3a7e3790a1
test(gate): establish the locale precondition instead of naming one
Review found the byte-versus-character witness asserting something
adjacent to its contract. It set `LC_ALL=C.UTF-8` and assumed the
locale took effect. Locale names beyond `C` and `POSIX` are
implementation-defined, so where that one is absent the shell falls
back to byte semantics --- and then the character-counting mutant
counts bytes too, agrees with the fix, and the row passes while
proving nothing. M-G-6 was killable here and unkillable elsewhere,
which is the same as not having it.

The locale is now chosen by BEHAVIOUR. Candidates come from `locale -a`
so the set reflects what is installed, and each is probed through the
same `/bin/sh` the gate runs under, asking `${#x}` on a two-byte
character and requiring `1`. No qualifying locale is a loud panic
naming what was tried, never a skip: a skip would be indistinguishable
from a pass, which is the failure mode this replaces.

M-G-8 proves the fix in three legs, because the hazard lives in the
environment rather than the code:

  8a  mutant gate, probed locale  -> the row fails, and the
      exact-boundary row still passes
  8b  SAME mutant gate, locale forced to `C` -> the row passes.
      The defect reproduced rather than argued.
  8c  no candidate can qualify -> panic naming the candidates

Also marks framing revision 6 approved and records M-G-8 in the ledger.

Gates: all nine green under `env -u TMPDIR`, log 20260813T182020Z.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 20:25:00 +02:00
Levi Neuwirth 0a04d55a35
fix(gate): review round 3 --- canonical ancestry, guard witnesses, and a withdrawn claim
**THE ANCESTOR WALK WAS WRONG TWICE OVER.** `for _anc in $(...)`
word-splits on IFS, so a gate root containing a SPACE was torn into
fragments and the real ancestor never tested --- the check passed on
exactly the path it should reject. And `dirname` walks LEXICAL
ancestry while `detect_project` canonicalizes, so a symlinked root hid
a marker the editor plainly sees. The walk resolves with `pwd -P` first
and iterates a quoted `while`; both shapes are verified by hand
(space-containing root refused, symlinked root refused at its real
path).

**THE 103-BYTE GUARD HAD NO WITNESS AT ALL** --- every other row runs
with a short root, so the guard is silent and a broken one looked
identical. Three rows now aim at it deliberately: boundary rejection
and acceptance, a MULTIBYTE root (each `é` is one character and two
bytes, so it is rejected only if the guard measures bytes), and
**rejection must reap both created areas**, which is the leak the early
trap exists to prevent.

**The `Cargo.toml`-DIRECTORY case was claimed and not covered**, and
the consequence is exactly as review predicted: reverting only the
language-marker arm to `[ -e ]` stayed green. The marker-type row now
drives all three shapes, and `M-G-5` --- that precise revert --- fails
it.

**Prose brought level with the implementation.** The framing, the
handoff and the ledger all said 108; the supported floor is **103
usable bytes**, Darwin's 104-byte array minus its NUL. The ledger also
still said `<pid>`, the superseded 21/30 reserve, and `M-G-1`.

**And the ruling said nested gates "do not pay" the reserve, which is
false and would have licensed exempting them.** They pay it in full;
the short layout merely gives them the headroom to satisfy an unchanged
production guard. Reworded, because the wrong version is the one a
future reader would act on.

**THE btrfs CAUSAL CLAIM IS WITHDRAWN.** The draft argued that a
one-second deadline plus a slower filesystem was a plausible new
mechanism for the fourth `managed_retry` occurrence. It does not
survive inspection: the deadline bounds the connection RETRY loop, not
the socketpair handshake that returned `BrokenPipe`, and the filesystem
work happens before it is armed --- the tempdir is created and never
bound. The environmental change is still recorded, as a CHANGE rather
than a mechanism, so a later occurrence can compare like with like.
Recording a mechanism the code does not support is worse than
recording none: the next occurrence gets measured against a story
instead of the evidence. TMPDIR stays disk-backed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 17:07:55 +02:00
Levi Neuwirth 1bd52b7f0d
fix(gate): review round 1 --- the propagation row proved nothing, and two guards were wrong
**THE PROPAGATION WITNESS DID NOT OBSERVE INHERITANCE.** The runner's
`eval` expanded `$TMPDIR` in the PARENT before `sh -c` ever started, so
the child received an already-substituted literal --- and an unexported
`TMPDIR=` would have passed the row unchanged. Single-quoted inside
`sh -c` now, so the CHILD expands it. **M-G-1b keeps the assignment and
removes only `export`: the row fails.** That is the mutation the
previous version could not catch, and the reason to prefer it over
M-G-1's blunter deletion.

**THE RESERVE WAS NOT THE MAXIMUM.**
`/.tmpXXXXXX/directory-target.sock` is 33 bytes
(`tests/gpu_invocation_acceptance.rs`), so paths of 76-78 passed the
30-byte guard and still blew the 108-byte limit during the CRDT sweep.
Reserve is 48 now --- the measured maximum plus ~45% headroom. And the
length is counted in BYTES: `${#var}` counts CHARACTERS under a UTF-8
locale while `sun_path` is byte-limited, so a multibyte path measured
short and passed a check it should fail.

**A MANAGED ROOT IS NOT INHERENTLY MARKER-FREE**, and assuming it was
rebuilt the original defect one directory up: a `.git` in `$HOME`, a
marker above `$HOME/build`, or a contaminated
`PMACS_GATE_TARGET_ROOT`. Placement under a directory the gate owns is
NECESSARY, NOT SUFFICIENT, and the old test proved only placement. The
gate now walks the ancestors and refuses, naming the marker it found.

`PMACS_GATE_ALLOW_ANCESTOR_MARKER` is the documented test-only escape,
beside `PMACS_GATE_TARGET_ROOT` in kind and risk: the behaviour tests
run under a tempdir whose ancestors they do not control, on a machine
whose `/tmp` carries this very marker, and their plans are synthetic so
no markerless fixture exists to re-root. **The check is witnessed by a
row that deliberately does not set it**, and M-G-3 (check removed)
fails that row.

**The guard leaked what it exists to manage.** It created both
temporary areas and exited before the trap was armed, so every
rejection left an AMBIENT and a TMPDIR behind. The trap is installed
first now; verified by rejecting a run and finding neither.

**`tmp/$$` with `mkdir -p` was not fresh.** PIDs are reused, so after a
SIGKILL it silently ADOPTS a leftover directory and the run inherits
another run's fixtures. `mktemp -d` fails rather than reuses.

**Prose corrected to match.** The handoff described
`<target>/gate-tmp/<stamp>-<pid>`; the implementation uses
`<gate-root>/tmp/<mktemp>`. Comments called the shared parent
per-worktree and pruned --- it is neither: `--prune` only considers
directories carrying an ownership marker, so the parent is skipped and
each run removes its own leaf.

**AND THE LANE CLAIMED A FRAMING EXCEPTION THAT DOES NOT EXIST.**
`AGENTS.md` says framing -> approval -> branch -> implement,
unconditionally; "the fix was already recorded as standing" is not an
exemption it grants. `docs/gate-script-framing.md` is amended as
**revision 6, AWAITING APPROVAL** --- a widening of §2's existing
isolation responsibility rather than a new feature, which is why it
amends that document instead of opening another. **This PR must not
merge before that revision is approved.**

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 16:34:21 +02:00
Levi Neuwirth 72647829df
docs: record PR #240 in the gate lane
The number goes in the moment the PR opens, per this file's own rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 16:00:17 +02:00
Levi Neuwirth d84df23aa7
fix(gate): isolate TMPDIR per invocation
Discharges the standing fix recorded in `docs/agent-handoff.md` §1 and
assigned to this lane. Every gate invocation now gets a fresh,
disk-backed `TMPDIR` at `<gate-root>/tmp/<pid>`, exported once so every
stage and every process they spawn inherits it, reaped by the same exit
trap as the ambient root. **A gate run no longer needs a `TMPDIR=`
override.**

**A CHILD OF `/tmp` WOULD NOT HAVE WORKED**, which is why the obvious
cheaper fix was not taken. The hazard is an ANCESTOR marker: project
detection walks upward, so a fresh subdirectory of `/tmp` inherits
`/tmp`'s ancestors and the same stray `.git`. The directory had to move
somewhere the gate already owns.

**`SUN_LEN` shaped the layout, and the fix's own gate run is what found
it.** A Unix socket path cannot exceed 108 bytes, and the suites bind
sockets INSIDE `TMPDIR`. The first placement --- `$TARGET/gate-tmp/$STAMP-$$`
--- produced a 114-byte socket path and failed SIX daemon and attach
tests with "path must be shorter than SUN_LEN". It hangs off the gate
root (36 bytes) rather than the per-worktree target (60) now, with a
short name: 47 bytes, leaving 61 for fixtures. Running the real gate
rather than only the witnesses is what caught this.

**A startup guard turns that failure class into a named one.** Six
socket failures deep in a suite name a LIMIT, not a CAUSE; the guard
fails immediately with the path, its length, and what to shorten.

**Its reserve is measured, not round, and the first value was wrong in
the more embarrassing direction.** The longest suffix a fixture appends
is `/.tmpXXXXXXX/test.sock`, 21 bytes, so 30 leaves ~40% headroom. An
earlier "generous" 45 FIRED ON THE GATE'S OWN BEHAVIOUR TESTS: they run
the gate inside the gate, so their root sits under the outer run's
TMPDIR and the nested path reaches 71 bytes. A guard that rejects a
legitimate configuration is worse than the failure it prevents, because
it fires on every run instead of a rare one. Verified both directions:
still catches an 87-byte root, silent on the real one and on the nested
tests.

**Two witnesses, each mutation-checked.** `M-G-1` removes the export ->
the propagation row alone; `M-G-2` stops the reaping -> the cleanup row
alone. Propagation is observed in a SPAWNED CHILD --- the self-test's
first step reports its own `$TMPDIR` into its log --- because asserting
the variable inside the script would only prove the script can set a
variable. The cleanup row runs against the self-test, which FAILS on
purpose, so it also pins that the trap fires on the failure path, which
is the path a leak would actually happen on.

One witness of mine needed correcting twice, both times because it
asserted something adjacent to the contract: first `!starts_with("/tmp/")`,
which tested where the FIXTURE put its root and failed on correct code;
then `contains("/gate-tmp/")`, stale after the directory was shortened.
It now asserts the exact parent, `<root>/tmp`.

**Proved against the live hazard:** `/tmp/.git` is still present on this
machine, and `m4_24_bare_string_glob_stays_relative` --- one of the two
tests it reddened --- passes with no override.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 15:55:19 +02:00
Levi Neuwirth 36a32961f2
docs: record PR #239 in the 1a lane
The number goes in the moment the PR opens, per this file's own rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-13 08:44:42 +02:00
Levi Neuwirth 69b9e6a41f
docs: record 1a's gate result and mutation inventory
Sixteen gates green with --protocol, under an isolated TMPDIR.

The crdt-gated wire suite is verified to have RUN --- 2 tests in the
crdt sweep, 0 in the default one --- rather than assumed, because a
suite that compiles to nothing reports ok and that is the shape this
project has been caught by before.

The mutation inventory names which row each one fails, including
M-1a-6: a mutation that targeted the wrong side of the append boundary
and so reported a sound pin as vacuous.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
2026-08-12 22:55:55 +02:00