Two corrections to U20 and the lane record for a2d5b26.
U20 said six green control runs establish that the observing diff is
not the cause. They do not. They establish non-reproduction in six
runs, which is all a rerun ever establishes in this registry --- a tree
that fails intermittently can carry a changed failure rate that six
runs are far too few to see. Treating non-reproduction as exoneration
is exactly the reasoning this file refuses when a rerun turns a red
green. What the controls actually do is remove the easy story and leave
the question open, and the row now says so.
The margin comparison was also arithmetic dressed as a phrase. "A third
again worse" than U6's 1.592 is not what 1.879 is: as a ratio it is
1.18x, as budget excess (0.879 over versus 0.592 over) it is 1.48x.
Both numbers are now given, with the note that they answer different
questions --- which is why neither gets compressed into an adjective.
The ledger records clause 5's replacement half and the two witness
repairs that came with it, both being assertions that looked strict and
were not: an origin that merely came down rather than landing on the
exact bound, and a rationale about the caret's position that the
fixture made false.
Three records, all from the same session.
**U20.** U6's composition-overhead test redded ALONE, its paired
keystroke test passing in the same run. U6's own closing rule says that
is a different incident, so it is filed as one rather than as a sixth
U6 occurrence: U6's selector requires the pair, and its whole argument
is that two unrelated subsystems failing at once is less likely than
one loaded machine. One test alone does not carry that argument.
The observing diff touches the paint path, so it was a live suspect and
was tested instead of argued about --- three full-lib runs with it and
three with the two files restored to HEAD, all six green. The margin is
recorded per U11: 1.879 against a 1.10 budget, a third worse than U6's
worst. The row says plainly that the size of the margin does not
resolve whether this is load or regression, and that a contemporaneous
load reading is the missing evidence.
**The shared-target false red fired a third time**, giving the complete
set of four error sites that the second occurrence's captured tail had
cut to three. An earlier draft of that entry guessed the missing fourth
was widest_display_columns; the third occurrence shows the guess was
right, and the entry now says it was still right not to record it --- a
signature that is usually right is one nobody can match against.
**And the latch entry was wrong.** It listed the manual horizontal
authority latch as "landed but not yet witnessed". It was written in
four places and read in none on the GPU, and absent entirely on the
TUI. An unread bool preserves nothing. The entry now carries the
measurement that showed it --- origin 30, next paint 0 --- and the three
things the framing's L-table did not anticipate, since those are what a
recovering session would otherwise rediscover from scratch.
The post-merge gate of the docs-only absorption commit redded at
`07-sweep` with U16's exact signature. Same selector, same panic site,
both required fragments.
It is the first occurrence on `main`, which removes the last attribution
question this row could have had --- the first two were on a branch
whose diff touched nothing under `src/packages/`, and this one has no
observing branch at all.
Three occurrences in about eleven hours, against eight consecutive green
`--lib` runs earlier the same morning. That is the closest this row has
to a rate and it is still not a measurement, because nobody has counted
runs and failures over a fixed window. Isolated reruns green three
times, which per the rerun rule establishes intermittence only.
The structural fix belongs to `file_io`: a version-suffixed child
inherits a cwd that a `TempDir` then deletes, and no care inside the
mutating test can close a process-wide window.
The heading said "a terminal bell never arrives within a 5s poll" while
the body two paragraphs down withdraws exactly that: the evidence shows
the bell was not OBSERVED within five seconds, not that it never came. A
title is the part most readers keep, so it was the worse place to leave
it. Retitled to match.
U16 said its reds and greens were "across two days". Every run the row
cites --- both reds and every green --- is 2026-08-31, which the
sentences immediately above it already said twice ("again the same day",
"returned within the day"). Corrected in both files.
U19 said "no rerun was performed on this selector". The very next gate
run was one: it passed in `03-lib`, `04-lib-crdt` and the exact
`07-sweep` context it failed in, at ea786a2. So did U16's selector,
after its second occurrence. Both rows now record those passes and
classify as intermittent. Writing "no rerun was performed" in the same
commit whose gate reran it is the kind of claim this file exists to
catch.
U19 also overstated three things. The evidence shows no bell was
OBSERVED within five seconds --- not that one "never comes", and not
that scheduling cannot explain it. "Three orders of magnitude of slack"
does not hold against the 200ms budget it cited: 5s is 5000x of 1ms but
only 25x of 200ms, so the distinction from the budget family is one of
degree. And adding an elapsed value later cannot make a future margin
comparable with THIS unmeasured one --- that margin is gone for good; it
only makes future failures comparable with each other.
R7 kept a sentence reconstructed before the renumbering: its "prior five
spread across lanes and months" were eight, and not evenly spread ---
three of them fall on one branch on 2026-08-15. The line census skipped
occurrence five, whose block never captured a line; it is now marked
unrecorded rather than guessed or omitted.
U16's introduction still said it was "worth more than its one
occurrence" while its status said second.
The gate verifying the R7 renumbering redded twice in `07-sweep`.
`cache_survives_across_fetcher_instances` is **U16's second
occurrence** --- same selector, same panic site, both required
fragments. It is the first time that row has reproduced, and it settles
an earlier withdrawal in the right direction: claiming "the window is
narrow" from eight green runs was wrong, and the failure came back
within the day. The child-inheritance chain stays a candidate; this
occurrence demonstrates it no more than the first did.
`terminal_bell_baseline_suppresses_history_and_delivers_each_new_bell_once`
is new, recorded as U19. It is a deadline but not the budget family's
kind: those assert work finishes in 1ms or 200ms, while this asserts an
event arrives at all inside FIVE SECONDS. Folding it into that family
would blur the one distinction those rows have.
Like R1, its assertion has `Instant::now()` in hand at the panic and
reports none of it, so the margin is unrecoverable and a future
occurrence will not be comparable to this one. That is the second place
in this codebase where the same omission costs the same thing.
`docs/active-work.md` recorded two full-fragment R7 occurrences from
2026-08-15 (logs 20260815T095532Z and T100719Z, `attach.rs:1728`) under
a heading saying they were "owed to the registry by whichever branch
merges second". Both branches merged. Nothing carried them across, and
they sat there for sixteen days, so R7's count read two low even after
yesterday's renumbering.
With those absorbed and the duplicate "fourth" fixed, the sequence is:
August 29 = ninth and tenth, August 30 = eleventh, August 31 = twelfth.
The parse-budget lane block, which still said sixth and seventh, is
updated too.
The deferral itself was reasonable --- this file has been bitten by two
branches inventing the same row id --- but not discharging it was not.
The lesson recorded is narrower than "absorb faster": an entry parked
under "owed to the registry" needs an owner named in the same sentence,
or it belongs to nobody.
The summary cell's line-specific claim was also incomplete: occurrences
one through FOUR report `attach.rs:1680`, not the first three.
And U18 over-corrected. `GONOSUMDB` is a real Go variable ---
`go help environment` documents `GOPRIVATE, GONOPROXY, GONOSUMDB` as
module prefixes "that should not be compared against the checksum
database", which is exactly the step that failed. It is technically
applicable; whether the authentication tradeoff is acceptable is a
different question. Discarding a real knob while correcting an invented
one is its own error and is recorded as one.
R7 carried TWO blocks numbered "fourth" --- D3 on 2026-08-11 and TMPDIR
isolation on 2026-08-13 --- so every later ordinal was one low. The new
red is R7's TENTH, not its ninth. Renumbered by date, with the duplicate
recorded in the status cell rather than silently fixed. The summary cell
said "three occurrences" and now states the total, keeping the
`attach.rs:1680` fact as the line-specific claim it always was.
U9's selector was named wrong in yesterday's correction. The row names
the CANONICAL pty test (`src/process.rs:3967`), not
`raw_mode_disables_kernel_echo` (`:3945`) --- and U9's own "relation to
U2" cell turns on exactly that distinction, so getting it backwards
would have undercut the row the paragraph was correcting.
The arithmetic was still wrong in two places. There are 121 result lines
and the doc-test groups are numbers 120 and 121, not 121 and 122; and
U9's table cell still asserted the strict 119-to-121 alternation that
the paragraph below it retracts.
U18 listed three outage options and all three were wrong. `GONOSUMCHECK`
is not a Go environment variable --- that sentence invented it.
`GOFLAGS=-mod=mod` selects module update mode and does not bypass
checksum-database authentication. And a version-suffixed `go install`
ignores vendor directories, so "a vendored gopls" needs a different
installation path. No replacement knob is named, because none was
verified.
The ledger called uncontrolled foreign load "the same evidence U9's
synthetic-load control was meant to produce" and said "U9 stays owed",
both of which contradict the correction below them. And its lane heading
said four registry rows moved while listing eight.
The gate verifying the previous commit redded at `gpu` with all three of
R7's required fragments, same selector and same line. Other seven stages
green; the observing commit is documentation only.
It adds a count and nothing else, which is the honest description. The
eighth occurrence's method note says the remaining candidates must be
varied inside the gate, one per run, and that is not this lane's work.
The loadavg reading is recorded as a condition, not a cause --- R7 is
not a budget row.
Three corrections to yesterday's correction, and one new row.
The serial-binary measurement was stated as "119 Running and 121 result
lines alternate strictly", which cannot be strict --- the count itself
gave it away. Precisely: 119 ordinary targets each report before the
next starts, and the two extra result lines belong to `Doc-tests pmacs`
and `Doc-tests pmacs_protocol`, which cargo labels differently and runs
last.
The replacement premise did not describe U9 either. Both U9 selectors
live in the ROOT LIB TARGET --- `m6_1_pty_raw_mode_disables_kernel_echo`
(src/process.rs:3945) and `composition_overhead_under_ten_percent`
(src/editor.rs:9717) --- and the sweep runs that target first, finishing
it in about 12 seconds. The sweep's later minutes cannot reach them.
What survives: the sweep re-runs the lib target late in the overall gate
invocation, under unmeasured machine state.
Two stale references to U9's void control are corrected, including the
ledger's claim that a synthetic-load run would "either implicate load or
clear it". It would not: with concurrency fixed at 1 there is no second
arm, so a red shows load is sufficient and a green shows nothing.
Non-reproduction never clears anything under this file's own rerun rule.
U18 is new and a new class. `Test (ubuntu-latest / luajit)` died in
toolchain setup before any cargo command ran: `go install gopls@v0.16.2`
hit an HTTP/2 INTERNAL_ERROR from sum.golang.org while verifying
x/telemetry. Every other row here is a test that failed; this is
infrastructure the workflow depends on failing to answer, and it
presents as a red check indistinguishable from a real one.
U9's "structural difference worth testing next" claimed that
`cargo test --workspace` runs many test binaries concurrently while
`--lib` runs one, and derived its discriminating control from that:
"pin test-binary concurrency to 1". The premise is false. Cargo runs
test TARGETS serially, one executable at a time, so that concurrency is
already 1 and the control pins nothing.
Measured in this project's own gate logs rather than asserted from the
cargo book: `20260831T093655Z-857818/07-sweep.log` alternates `Running`
and `test result:` strictly --- 119 to 121 markers, ZERO cases of one
binary starting before the previous reported. The pattern is `RTRTRT`.
That falsifies a premise two rows rested on, so U12's family paragraph
is corrected too. What survives is smaller and still true: a sweep is a
long sequence of binaries, so a budget inside it runs at an arbitrary
point in a multi-minute step. The family still should not consume review
rounds --- but it now needs a control someone has to design.
U17 no longer claims `--test-threads=1` exercises U9's control. It is a
different knob: it serializes test FUNCTIONS within one executable. Its
candidate mechanism is narrowed to match --- removing sibling test
functions removes ONE source of contention, which supports neither
"fastest" nor "narrowest".
R6's block drops two overclaims: a PR run CAN show the identical red
(only the main dispatch establishes it on the merge base), and this was
not the dispatch key's first use --- #245's D2/D3 dispatched three runs
right after it merged. It is the first use for a live merge-base
control.
Five corrections, all mine.
U16 stopped at "the window exists", which misses why restoring the cwd
does not close it. `run_git` calls `run_git_inner(None, ...)`, and that
sets `current_dir` only when `cwd` is `Some` (fetcher.rs:329-330), so
the spawned git INHERITS the parent's temporary cwd. The parent then
restores its own --- which does nothing for a child that already has its
working directory --- and the TempDir drops underneath it. The restore
is not merely too early; it is irrelevant to the child.
U16 also offered a serial guard around `set_current_dir` tests as a
structural control. That does not protect an unguarded test that spawns
a child, because the child outlives the guard. The options that work are
removing the cwd mutation, running that test in a subprocess, or
serializing the whole lib-test binary.
And U16 said 8 green runs showed the window was narrow. They do not.
Non-reproduction establishes intermittence and nothing else; nothing
here has sized this candidate's window.
U17 claimed no PR run can show its failure. A PR run exercises the same
test and could fail identically; what only a main-side run establishes
is that it fails ON MAIN, with no observing branch to suspect. And its
`got ok` does not prove the supersede arrived late --- it proves the
predecessor completed successfully before cancellation took effect,
which a timely supersede whose cancellation lost the race produces
identically.
The macOS lua54 leg redded on PR #246 with
`acc28_child_input_and_the_c_c_escape_work_unchanged_in_a_panel`. It is
a full three-condition match for R6 --- selector, flavor, and BOTH
required fragments (`timed out waiting for` + `/ready`) --- 26 days
after the first occurrence.
The log was read BEFORE anything was rerun. U3 named that lesson and U8
recorded its fourth violation; this is the first time it was followed on
a macOS job at the moment it mattered, and the fragments exist because
of it.
Rather than argue from an unrelated diff, a merge-base control was
dispatched at `aae5b35` --- the first real use of the
`workflow_dispatch` key #245 landed, and exactly the case U11 motivated
it for. The macOS legs came back GREEN, so the inference the control
could have supplied is unavailable. Recorded as a null result, the way
R1's row had to record its own. What each outcome would mean was written
down before the result was seen.
The control was not otherwise clean: `Test (crdt)` failed on `main`,
which is U17. It fails the opposite way to R1 and R5 --- not a missed
deadline but a predecessor that had already completed --- and the job
runs `--test-threads=1`, the condition U9's still-unrun control names. A
red on the merge base is invisible to any PR run.
The sweep step redded on `cache_survives_across_fetcher_instances` with
`fatal: Unable to read current working directory`. Not a budget test,
and not a load story.
It is the only row in this file that arrives with a named candidate
mechanism inside the test suite. `src/file_io.rs:434` calls
`std::env::set_current_dir` --- process-global state --- inside a test
running in one of libtest's parallel threads, points it at a `TempDir`,
and lets that `TempDir` drop. Every other test in the binary shares that
cwd for the window, and after the drop it is a deleted directory, which
is exactly what git reported.
Recorded as a candidate with a citation, not a demonstrated chain: 8
full parallel `--lib` runs did not reproduce it, which says the window is
narrow rather than absent. The row names the two controls that would
settle it and runs neither --- the structural fix is `file_io`'s, not a
CRDT invariant lane's.
Two precision errors, both mine.
U14 said "four selectors in three unrelated subsystems". They are four:
the async runtime, the optimistic-echo orchestrator, editor composition,
and the LSP dispatch seam. U6's own row treats its two selectors as
unrelated subsystems, so the `04-lib-crdt` pair is two of the four here,
not one. The `what is NOT` row said three as well.
U15 claimed to be the load number "U6 and U7 have each wanted since
2026-08-09". Half of that was wrong: U7 has carried a load average
(12.9 / 23.9) in its job/flavor field since that date. What U7 records
as unmeasured is narrower --- whether the shared `CARGO_TARGET_DIR` and
its sibling builds PRODUCED that load. So 34.04 is the first
contemporaneous reading for a U6 occurrence and a second data point
beside U7's, not the registry's first. The row title oversold it too.
U15's disposition list also omitted U15.
Three corrections, all mine, all caught in review.
U14 claimed a second occurrence for a run whose selector set had
ROTATED --- `full_buffer_summary_flatten` and `dired_renders_10k_entries`
in place of `grep_supersede` and `acc34_purge`. This file's own matching
rule requires the exact selectors to match, so that is a new incident.
It is now U15. The `04-lib-crdt` pair the two runs share is recorded
where it belongs, as U6's own occurrence; U6 goes from one occurrence to
five, four of them on 2026-08-30.
U14 also said "three unrelated tests". There are four selectors.
And the load claim went too far. `/proc/loadavg` was read once, after
the second run, so there is no series to correlate against; the margins
are not monotonic (`composition_overhead` ran 1.182x, 1.592x, 1.527x,
and reports two different values within the second run); and an earlier
version said a load average of 34 "explains it without any help". What
34.04 establishes is severe unrelated load present CONTEMPORANEOUSLY
with one multi-red run --- a measured confound, not a measured cause.
That is still worth more than U6 and U7 have had since August, and it is
worth exactly that much.
The lane block also still carried the withdrawn "opposite way to R7"
claim and named `db24ae3` as the gate head. It now names `2c24303` and
log 20260830T193305Z-4167110.
The gate redded again in the same three stages, with a partly rotated
selector set --- one of them being a U7 selector. This time
`/proc/loadavg` was read at the failure: 34.04, with the CPU saturated
by an unrelated `lean` workload on this shared machine and no cargo,
rustc or gate process of mine left running.
U6 and U7 have each recorded, since 2026-08-09, that the load confound
"was not measured, so it is a rival explanation, not a finding." It is
measured now, and the margins move with it monotonically across three
runs of one unchanged tree: 1.343883ms, then 1.689259ms, then
2.269247ms, against a 1ms budget. A regression does not get 69% worse
between two runs of the same tree.
This retires nothing. The budgets are still wall-clock assertions whose
measurement design nobody has defended --- R1's disposition, applied to
five more tests. What changes is that "one loaded machine" is now a
measured explanation rather than a plausible one.
The gate run verifying revision 5 redded three unrelated tests in three
stages: a 50ms supersede budget in `lib`, U6's pair in `lib-crdt`, and
an LSP readiness race in `sweep`. Recorded as U14, because the
co-occurrence is the signature --- three subsystems failing in one run
is far less likely than one loaded machine, and no selector reds twice.
It also falsifies something I had committed an hour earlier. U6's
second-occurrence block said the row "runs the OPPOSITE way to R7",
resting on both failures being out of gate while `04-lib-crdt` was green
in four gate runs. The next gate run redded `04-lib-crdt` with exactly
that pair. Four green stages were a run of four, not a property. The
claim is withdrawn in place rather than edited away, and U6's status
moves to a third occurrence: three in one afternoon, twice out of gate
and once in.
U14 also declines an R1 match it could have claimed. The `lib` failure
carries R1's required fragment but a different selector, and this
registry matches on both. Worth noting separately: the sibling test
already reports the elapsed value R1's row records as missing from its
own assertion --- the cheap half of what R1 defers is written next door.
R7's eighth-occurrence write-up said the gate's ambient root, TMPDIR and
cross-stage process state were "now the only place the difference can
be". That is wrong. The paired runs exclude the SOURCE TREE and nothing
else: scheduler load, kernel and socket timing, page cache pressure and
whatever else the machine was doing also varied between them, and a
BrokenPipe on a socket handshake is exactly what those can drive. The
three remain the candidates worth varying one at a time --- because they
are the ones this project can vary --- not an exhaustive causal set.
U6 gained a second occurrence, and for the first time it REPRODUCED:
both selectors, both fragments, two consecutive runs. Margins recorded
per U11's lesson --- 1.343883ms against 1ms, and 1.182x against 1.10x.
Its asymmetry runs the opposite way to R7's: both failures were out of
gate, while the same command as `04-lib-crdt` was green in all four of
this lane's gate runs. Whatever the two rows share, it is not a
direction.
Framing revision 5 and the lane block are updated to match, including
the stale "AWAITING APPROVAL. Nothing implemented." header and the gate
line that named a commit the branch had already moved past.
Two consecutive gate runs on one worktree, minutes apart. Heads differ
by a single commit touching a single markdown file. The first was all
eight stages green; the second redded at `gpu` with all three of R7's
required fragments, and at `sweep` with the same single test.
The fifth occurrence excluded the observing tree relative to `main` by
having a documentation-only diff. This pair excludes it relative to the
immediately preceding GREEN RUN OF THE SAME GATE on the same worktree,
which is strictly sharper --- whatever varies across that green/red
boundary, it is not the source tree.
No ratio is claimed from it. Folding verification gates into the
2026-08-29 window is exactly the drift that window was bounded against.
Five isolated selector runs were green, which per this file's own rerun
rule and the seventh occurrence's correction establishes intermittence
and excludes nothing.
Recorded together, because the control is what the occurrence owed.
The occurrence: Test (macos-latest / lua54) on head d9cc0fa, 1984 passed
and 1 failed, panicking at src/async_runtime.rs:2444 with R1's required
fragment. THIS IS R1 despite the new Lua flavor --- the row records
luajit and this is lua54, and flavor is not part of signature matching.
The selector and the fragment are, and both match.
The control: the same job at the branch's EXACT merge base 2e9f62b,
rerun the same hour, green at 1985 passed and 0 failed. Both logs were
preserved before the rerun, so neither result is reconstructed from a
conclusion --- the selector is read as "... ok" in the control's own log.
What that green control establishes is carefully bounded. It does NOT
establish environmental cause and does NOT retire R1; it shows only that
the merge base can pass the same job in the same hour. A RED control
would have established that the branch did not introduce the occurrence,
and that inference is simply unavailable here. R1 remains live either
way under its measurement-design disposition.
Records that R1's assertion omits its measurement --- the same class #244
fixed, and a line my discarded sweep had surfaced --- and that it is
deliberately NOT fixed here. Adding the elapsed value would sharpen the
next failure's evidence and repair nothing about the measurement design
this row is about: that a thread::sleep(15ms) is asserted by comment to
mean the worker picked the job up. That belongs to the async-runtime
lane with the rest of Q#MCI3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The head-exact review gate reached sweep with six green stages, then
`skipped_directories_are_reported_with_a_reason` observed empty child
stdout. The durable failure cannot say why because the row discards the
child status and stderr.
Record the exact signature, the structurally unrelated branch diff,
and the isolated-selector and full-binary green reruns at their actual
strength: intermittence only. Diagnostic hardening remains a separate
lane.
The four-run observation window contains two green in-gate runs, so it
cannot be described as ending at the first green. Name it directly as
the first four in-gate runs; the timestamps continue to define the
boundary exactly and later verification gates remain excluded.
The previous commit recorded "in-gate 2 failures in 4 runs" and named
the final verification gate. Re-running the gate on that very commit
made both wrong: a fifth in-gate run, green, and a new log id. Every
docs fix forces a re-gate, and the re-gate invalidates the docs fix.
R7's window is now explicitly the four in-gate runs of 2026-08-29 up to
and including the first green, plus the 17 out-of-gate runs taken
between them. Later verification gates are excluded by definition. A
ratio that grows with review activity measures review activity, not the
phenomenon.
The ledger stops naming the head and log id at all. It records that
every review round ends with a green head-exact gate and points at the
PR body for the current values --- which is the same reason SS5b stopped
recording an ahead-count: a line naming a moving value is stale the
moment it is written, and that lane learned it twice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Four stale claims, all in durable documentation that would have shipped
as written.
The framing still said "AWAITING APPROVAL. Nothing implemented" while
the ledger and the PR both said otherwise.
R7's status line said SIXTH OCCURRENCE while the body below it recorded
the seventh. Its tree-exclusion bullet said the observing lane's diff is
two docs; the lane carries three. And its cumulative figure said in-gate
is 2 failures in 3 runs, which omitted this lane's final head-exact
verification gate --- a fourth in-gate run, green. The figure is 2 in 4,
and the fourth is named so a reader can tell which run it was.
Active-work carried the same stale docs count and the same 2-in-3, and
recorded neither PR #244, nor the final head 2d76984, nor the gate log
20260829T152824Z-673477. It now carries all three, with the gate's
result read from its eight stage logs rather than inferred from stage
exits.
The sweep that found the last 2-in-3 also printed "(none = consistent)"
unconditionally, which is the read-success-from-absence shape this
project keeps catching. It now prints the hit count and only claims
CLEAN when that count is zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The gate run immediately after the seventh occurrence passed all eight
stages with zero failures anywhere, R7's selector included. So "in-gate
always fails" is false, and the previous entry --- written before that
run --- is corrected rather than deleted.
The correction that matters is a reasoning error, not a data one. I
called seventeen green out-of-gate runs "four hypotheses excluded". They
exclude nothing: NOTHING outside the gate has ever reproduced this
failure, so matching one gate condition at a time outside the gate
cannot isolate an in-gate cause. All those runs establish is that none
of the four conditions reproduces it BY ITSELF.
Stated at the strength it carries: in-gate is 2 failures in 3 runs,
out-of-gate is 0 in 17. Suggestive, and not a clean split.
The method follows from that and is written down for the next
occurrence: varying conditions outside the gate cannot answer this
question, so the gate's ambient root, its exported environment, and
process state across stage boundaries each need a gate run with that one
thing changed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The next gate run of the same worktree reproduced R7 immediately --- same
selector, same three fragments --- and that run's other seven stages were
green, sweep included and complete, so the pair is not confounded by the
truncation that marred the sixth.
Two consecutive in-gate failures is new for a row whose prior five were
spread across lanes and months, so it prompted a narrowing. Seventeen
green runs at the failing head on the failing worktree exclude four
hypotheses: the selector being flaky (3 isolated), the gpu binary's own
concurrency (6 full runs), the gate's isolated TMPDIR (6 runs under a
gate-shaped 61-character path, tested because this project already knows
socket-path length matters), and residue from the m4 stage the gate runs
immediately before gpu (2 back-to-back pairs).
So the discriminator is inside scripts/gate versus outside it, and it is
not the TMPDIR, not the preceding stage, and not the binary's
concurrency. The row does not guess at what remains --- the gate's
ambient root, its exported environment, and process state carried across
stage boundaries are named as uneliminated, not as suspects.
Causal status stays UNRESOLVED, but the question is sharper than it was:
previous entries compared trees and lanes, and this one locates the
difference in the runner.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The gpu stage failed with all three of R7's required fragments ---
`transient sequence must attach`, `Handshake(Io(`, `BrokenPipe` --- at
pmacs-gpu/src/attach.rs:1889, 283 passed and 1 failed. Isolated selector
green three times afterwards, which is this row's established control.
The tree exclusion is as strong as the fifth occurrence's: this lane's
entire diff is src/async_runtime.rs, tests/m4_acceptance.rs and two
docs. No pmacs-gpu file is touched, and the change is two assert!
message strings.
Records the run's other half honestly. The gate was in a background task
killed at 314s and 07-sweep.log ends in Terminated, so the run is NOT a
gate result. Stage 6 is evidence because it completed and reported;
stage 7's absence is evidence of nothing. Distinguishing those is the
same rule the handoff already carries about timeout-wrapped gates.
Causal status unchanged: UNRESOLVED. A sixth lane touching a sixth
unrelated surface saw the same three fragments, which strengthens
"not lane-correlated" and settles nothing else.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Head-exact gate at 6142acc: fifteen of sixteen stages green, including
sweep, m4, gpu, diff-check and all eight touched acceptance suites. The
red was 04-lib-crdt, where composition_overhead_under_ten_percent
(1.247x against a 1.10x budget) and
setsid_escapee_is_not_reaped_and_teardown_reclaims_readers failed
together. Both green on isolated rerun.
Filed as U12 rather than folded into U6 or U9, because both of those
instruct it: U6 says one of its selectors redding without the other is a
separate incident, and this is composition_overhead alone for the second
time; U9 is the same budget-plus-PTY shape but in 11-sweep with a
different PTY selector.
src/process.rs is not touched by this branch at all. src/editor.rs is,
but only in the panel-replay paths, not in composition.
The row does NOT claim load caused it. It records that the run was
knowingly taken on a machine that was quieter but not quiet --- load
11.04 at the start, 27.79 five-minute at the end, two foreign python
processes throughout, an apt install shortly before --- which are
conditions, not a mechanism.
Four incidents in this family now, and the discriminating control U9
named remains unrun: pin test-binary concurrency to 1, and separately
load a lone --lib binary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
SS5b merged as #242 (47b5463) at approved head 61f0faf, via
--match-head-commit so the merge is provably of the reviewed head.
PROTOCOL_VERSION is now 25. panel-pointer-replay is unblocked and is the
next step in the arc.
#239 (ca92796) and #240 (72da24a) both merged on 2026-08-13 and the
ledger has been calling them OPEN for a week. Their blocks are kept for
their reasoning, relabelled for their status.
#240's block gains the postscript it earned: its TMPDIR isolation was
the thing I defeated during SS5b review round 4 by running the CRDT
sweep by hand, outside scripts/gate. Two m4_24 base-resolution rows
failed, I reported them as pre-existing and proposed CI as the arbiter,
and the actual cause was /tmp/.git being inherited as an ancestor
project root. Through the gate, both pass.
Adds U11 to the red registry, the row deferred during #242's review so
that no docs commit would invalidate that PR's head-exact gate evidence.
It carries the exact selector and panic fragment, both attempt IDs, the
1960/1 counts, and the fact that the margin is unrecoverable because
duration_ms is omitted from the assertion message --- which is why a
recurrence owes a merge-base control rather than a comparison. The four
exact-head local passes, the two macOS/luajit greens and the identical
async_runtime.rs blob are recorded as narrowing evidence and explicitly
not as causality.
Per the standing rule, this absorption does not advance any canonical
base to its own commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Two consecutive scripts/gate runs at 70b334d, worktree verified clean
before and after each run, were 15/16 green apiece. Run A redded
13-sweep on dired_open_renders_10k_entries_under_200ms at 263.961465ms
against a 200ms budget; run B redded 15-sweep-crdt on
criterion_1_end_of_line_typing_completes_sub_frame_per_keystroke at
1.044609ms against a 1ms budget. Each red is green in the other run,
and both are green isolated at load 9.34.
That excludes the tree more strongly than U7 could: not "the diff
touches no render path" but the SAME COMMIT passing and failing each
row. Neither failing path is touched by the branch under test.
It does NOT establish load as the cause --- load was not sampled during
either failing step, and the row says so rather than borrowing a
reading taken elsewhere in the run.
Honours both escalation rules it trips. U7 says a repeat of one of its
selectors is a separate incident, and run A repeated one; U6 says one of
its pair redding alone is a separate incident, and run B did that. Both
are filed here rather than appended. The row also declines to pick
between the repeat and the rotation, because both are true of these
observations.
Closes one rival U7 left open: per-worktree gate target directories mean
no sibling shared this one.
Names the standing discriminating control U9 already specified and which
remains unrun --- pin test-binary concurrency to 1, and separately load a
lone --lib binary --- and asks that this family stop consuming review
rounds until it runs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Answers review of 15. Framing only. Three items reverse a rule 15
introduced, and one retracts a mutation that was not a defect.
**THE PRODUCER RULE CONTRADICTED PROACTIVE CANCELLATION.** The daemon
cancels BEFORE emitting the replacement frame, so the frontend needs
only to clear its local latch when that frame arrives, and send
nothing. Revision 15 asked it to emit a cancellation tail or retain the
latch: the tail is redundant --- the daemon would receive a release for
a gesture it has already settled, which is the duplicate release the
latch exists to prevent --- and RETAINING IS ACTIVELY HARMFUL, because
it manufactures a `Drag` under the NEW generation with no accepted
`Down`. That is the exact orphan the section exists to prevent,
produced by the rule meant to prevent it.
Ordering is what makes the simple rule safe: cancel, then emit. The
frame's arrival IS the cancellation signal; no second channel is
needed. Witnessed as `Down` -> key advances -> replacement frame ->
motion and physical `Up` produce no new drag and no duplicate release.
**THE LATCH HAD ONE TRIGGER AND NEEDED FIVE.** Cancellation runs on
every loss of gesture authority: generation advance, `Absent`, panel or
buffer identity change, geometry-epoch change EVEN AT AN UNCHANGED CELL
TOTAL, and detach. And an ordinary accepted `Up` must clear the latch,
or a later invalidation finds a gesture it believes live and
synthesises a duplicate release for a button already up --- the replay
lane's D1/D2 orphan race, arriving from the daemon's side.
**G9b's MUTATION WAS A VALID IMPLEMENTATION, NOT A DEFECT.** Keying the
dedupe by `(mapping_generation, coord)` preserves same-generation
suppression and naturally admits the first motion under a new
generation. Requiring it to fail would have forbidden a correct design.
Replaced with two real defects: compare only the cell and never key or
reset by generation (the first post-change motion is eaten), and reset
on every same-generation repaint (pixel-rate traffic returns).
**"PROJECTED CELL IDENTITY" CONTRADICTED THE STYLING CONTROL** in the
same section. The wire `Cell` derives `PartialEq` over `glyph`, STYLE
and `attachment` (`pmacs-protocol/src/cell.rs:153`), so an identity
keyed on cell equality moves on a pure recolour --- while the stable
controls rule style out. Terminal identity is now glyph and row
TOPOLOGY plus the view anchor, excluding face, style and cursor, with a
same-glyph/different-style control: the row that catches an
implementation reaching for `Cell` equality because it is right there.
P2s: zero-generation rows added in BOTH directions as independent legs
(a valid `PresentMapped` with generation zero must be rejected
atomically; a zero-generation `PanelPointerMapped` must be refused);
G7 split into outbound mapped-frame and inbound mapped-pointer legs,
since its old mutation only withheld the frame; G2's grid rows/columns
and fold-map-content/`fold_projection`-policy composites split; and
SS20 now names journey steps 5 and 8 while stating neither grade
changes --- an auditor scanning for grade movement alone would
otherwise conclude this slice touches no journey.
**AND R7 RECURRED, ON A DIFF THAT IS ENTIRELY DOCUMENTATION.** The
first `--protocol` run of this tree failed the `gpu` step on
`managed_retry_survives_transients_and_uses_the_successful_stream`,
with all three required fragments verified from the durable log
(`20260815T072601Z`). Recorded as R7's FIFTH occurrence.
It carries the strongest tree exclusion the row has had: occurrences 1
and 4 argued "unrelated lane", while this branch cannot be related at
all --- no Rust, no wire surface, no `pmacs-gpu` file. The line moved
to `attach.rs:1728` from `:1680`, which the row already treats as
occurrence-specific rather than a fragment. Isolated rerun green, and
the full gate green on the re-run (271/271 in the `gpu` step) --- which
per this file's rerun rule establishes INTERMITTENCE ONLY, though here
there is no tree change to exonerate.
What five occurrences across three flavors and five unrelated lanes now
support is that the failure is NOT LANE-CORRELATED. That is evidence
about where the cause is not. The retirement condition is unchanged.
Gates: all eleven green under `env -u TMPDIR` with `--protocol`,
log 20260815T073556Z.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
**THE ANCESTOR WALK WAS WRONG TWICE OVER.** `for _anc in $(...)`
word-splits on IFS, so a gate root containing a SPACE was torn into
fragments and the real ancestor never tested --- the check passed on
exactly the path it should reject. And `dirname` walks LEXICAL
ancestry while `detect_project` canonicalizes, so a symlinked root hid
a marker the editor plainly sees. The walk resolves with `pwd -P` first
and iterates a quoted `while`; both shapes are verified by hand
(space-containing root refused, symlinked root refused at its real
path).
**THE 103-BYTE GUARD HAD NO WITNESS AT ALL** --- every other row runs
with a short root, so the guard is silent and a broken one looked
identical. Three rows now aim at it deliberately: boundary rejection
and acceptance, a MULTIBYTE root (each `é` is one character and two
bytes, so it is rejected only if the guard measures bytes), and
**rejection must reap both created areas**, which is the leak the early
trap exists to prevent.
**The `Cargo.toml`-DIRECTORY case was claimed and not covered**, and
the consequence is exactly as review predicted: reverting only the
language-marker arm to `[ -e ]` stayed green. The marker-type row now
drives all three shapes, and `M-G-5` --- that precise revert --- fails
it.
**Prose brought level with the implementation.** The framing, the
handoff and the ledger all said 108; the supported floor is **103
usable bytes**, Darwin's 104-byte array minus its NUL. The ledger also
still said `<pid>`, the superseded 21/30 reserve, and `M-G-1`.
**And the ruling said nested gates "do not pay" the reserve, which is
false and would have licensed exempting them.** They pay it in full;
the short layout merely gives them the headroom to satisfy an unchanged
production guard. Reworded, because the wrong version is the one a
future reader would act on.
**THE btrfs CAUSAL CLAIM IS WITHDRAWN.** The draft argued that a
one-second deadline plus a slower filesystem was a plausible new
mechanism for the fourth `managed_retry` occurrence. It does not
survive inspection: the deadline bounds the connection RETRY loop, not
the socketpair handshake that returned `BrokenPipe`, and the filesystem
work happens before it is armed --- the tempdir is created and never
bound. The environmental change is still recorded, as a CHANGE rather
than a mechanism, so a later occurrence can compare like with like.
Recording a mechanism the code does not support is worse than
recording none: the next occurrence gets measured against a story
instead of the evidence. TMPDIR stays disk-backed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
**The gate's own run reproduced a REGISTERED signature**, and it is
recorded as a fourth occurrence rather than waved through: same
selector, same `gpu`-step flavor, all three required fragments verified
against the durable log. Three isolated re-runs were green, which this
file's rule says establishes intermittence only.
**This lane is code-neutral for `pmacs-gpu` but NOT
environment-neutral**, and that distinction is the entry's point.
Occurrence 3 excluded "the added GPU test is the mechanism"; this
occurrence adds nothing to that binary at all, which corroborates the
exclusion independently. But the lane moves `TMPDIR` off `/tmp`, taking
every `tempfile::tempdir()` in the run from **tmpfs to btrfs** --- and
the failing test runs a handshake against a **one-second deadline**. A
slower filesystem under a timing-bounded test is a plausible mechanism
that did not exist in occurrences 1-3. Booking this as "the usual
flake" when the observing lane changed the conditions the flake is
sensitive to is exactly the reasoning this registry exists to prevent.
Also: the suite's own roots move to a short base. Rooting them under
the ambient `TMPDIR` put a NESTED gate's TMPDIR near 70 bytes, which
legitimately tripped its own SUN_LEN guard --- the suite failing on a
configuration it created rather than on the behaviour under test. And
the marker row's `.then(..).unwrap_or_else(..)` chain is gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Rebased onto e67ad07; the framing commit replayed with no conflict.
Three corrections, each verified against the tree rather than against
this lane's own notes:
COHERENCE.md said `v6..=v21` schema support. The ceiling moved TWICE
since: #221 to v22 for `LineWrapFacts`, #228 to v23 for
`MinibufferPromptRows`. Now v23. The same row's "production attach
remains v20" is correct --- `ADVERTISED_PROTOCOL_VERSION` is 20 --- and
is deliberately left alone.
Journey step 11 read "Works but undiscoverable --- no statusline
spinner/progress indicator anywhere". #232 shipped exactly that
indicator on 2026-08-09, so the row went stale the day it landed. Now
Partial, with what actually exists and what does not: the indicator is
there, the `*workers*` view still has no keybinding.
§9's "No progress indicator exists anywhere" carried the grep that was
the evidence for opening worker identity in the first place. Corrected,
with the part that did NOT change stated as plainly: a purpose says
what a job is doing, never who asked, and attribution is what §9
grades. THE SECTION'S GRADE IS LEFT UNTOUCHED pending a re-audit ---
moving a grade is an audit act, not a documentation correction, and
Stage 0 is docs-only.
U9's row claimed "whatever this is, it is not the tree" about a
same-tree green. This file's own rerun rule forbids that: a same-tree
green establishes intermittence only, and a tree can raise an
intermittent failure RATE without making it deterministic. Replaced
with "not deterministic on this tree; causation and rate effect
unresolved", and the wrong claim is quoted rather than deleted, because
it is the one a later reader would otherwise reach for.
Also corrects this lane's own earlier claim that `add0ba1` had done
half of Stage 0's absorption. It absorbed #227 and #234; five stale
lanes and both COHERENCE corrections remained.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Review round three on PR #235, plus the diagnosis of its first CI run:
all five test legs failed on one test, deterministically, while
sixteen local cores stayed green.
The mid-walk cancellation bound could not bite the per-entry poll
alone. The cancel lands two files into a 41-file directory --- root
contributes three dir entries --- so with the per-entry poll deleted
the directory finishes and the per-DIRECTORY poll catches at seen ==
44, under the old bound of 60. The bound is now 40 against an expected
exactly-35 (3 + one 32-entry poll stride), and the entry-poll-only
bite goes red at 44. Verified both ways.
The retirement helper observed a REQUEST, not settlement: it returned
as soon as an active row showed cancel_requested, which a worker that
ignored the token and completed successfully would satisfy. It now
waits for a completed row with status == "cancelled", making the
lane's "settles cancelled" claim true at the witness, not just at the
Rust layer.
The CI red: d3_pump(1600) between the mid-walk join and the late.bbb
write assumed the held walk would complete within 1.6 s. On a 3-thread
CI pool, 8 sleeps of 1200 ms drain in ~3.6 s of waves, so the file
landed before the held walk even STARTED and folded into the joiner's
baseline --- exactly the fold the test exists to assert for mid.bbb,
applied to the wrong file. Deterministic on every 2-4-core runner,
invisible on 16 cores. The drain is now an observable condition ---
at least one post-join walk completed and none active --- with the
saturation sleeps at 800 ms, and the three saturation tests plus the
whole eighteen-test family re-run green under taskset -c 0-3, the CI
pool shape reproduced locally.
A fixed-duration pump against pool-dependent timing is a core-count
assumption in disguise; the lane records it as such.
Superseded round-one text in the lane (the fallback "unreachability"
claim round two disproved) is corrected in place.
One gate run also hit the live attach-retry BrokenPipe row --- fourth
occurrence, all three required fragments verified against the durable
sweep log, recorded in docs/ci-red-signatures.md. This lane touches no
pmacs-gpu code, no wire, and no protocol; the same sweep passed twice
earlier the same day on materially the same tree. The retirement bar
(mechanism, not rate) is unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The merge gate for this lane failed at `11-sweep` on two selectors, and
both had already passed in `03-lib` and `04-lib-crdt` of the SAME gate
invocation, minutes earlier, on the same tree and machine. U6 and U7
could only ever compare a red run against a different run; this is the
first occurrence in the family where the control is inside the run, and
that is what the row is for.
Both are near misses against existing rows, and neither is folded in:
- The PTY failure carries U2's exact fragment, but U2's selector field
names only the *raw* selector. U2's occurrence 2 had raw and canonical
failing together; here canonical redded ALONE and raw passed, which
U2's evidence has never shown.
- `composition_overhead_under_ten_percent` is one of U6's two selectors,
and U6 instructs in its own text that one-without-the-other is a
different incident. It redded without its pair, in a different step,
at 1.613x against U6's 1.297x. Judged as instructed.
The row also records the first checkable candidate this family has had.
`cargo test --workspace` runs many test binaries concurrently while
`--lib` runs one, so the passing and failing steps differ in kind and
not merely in load average --- with a stated control that separates
load from concurrency. U6 and U7 both left the confound atmospheric
and unmeasured; this does not measure it either, but it names something
that can be.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Merged rather than rebased. Eighteen commits replayed against a ledger
that three other lanes had rewritten meant eighteen conflict
resolutions in `docs/active-work.md`, each one a chance to lose a lane
entry; merging resolves it once, against the state that actually ships,
and leaves the reviewed commits' SHAs intact. Only one file conflicted.
`docs/ci-red-signatures.md` auto-merged **without a conflict** — the
same silent path that produced duplicate U4/U5 ids when #232 rebased.
Verified by hand afterwards: ids U1-U8 are disjoint. They are out of
numeric order (U6/U7 sit ahead of U4/U5) and are left that way rather
than moved, since the note at the U6 row explains the history and
relocating sixty lines inside a merge commit hides real changes.
Three leftover conflict markers were sitting in `docs/active-work.md`
on `main`, committed by an earlier lane's resolution. `git diff --check`
flags them — but only for a working-tree diff, which is why the gate's
`diff-check` step never saw them and they survived several merges.
Removed here.
The U4 row is corrected on evidence this lane produced:
- **Flavour was wrong as a matching key.** The row was filed from
#229's `lua54` red and put the flavour in the key; #231 reddened the
identical selector with the identical three fragments twice on
`luajit`. Matching as filed would have missed both.
- **A fourth sighting was a deliberate bite, not an occurrence** — the
defect reintroduced on purpose during the test's own development. It
is recorded for what it proves instead: the genuine defect and these
CI reds are signature-indistinguishable, same message class and same
full-timeout duration.
- **The control experiment is written down with its own bounds** — five
green base observations against 0/2, 4.8% under an equal-rate model,
and the two facts that bound it: attempt 5 reddened a different
selector, and the branch side was never resampled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Attempt 5 of the merge-base control at 0190102 failed on
acc28_child_input_and_the_c_c_escape_work_unchanged_in_a_panel, a
selector in no registry row. I then reran that job before reading its
log, and GitHub keeps only the latest attempt logs for a rerun job, so
the assertion text is gone. Recovery was attempted through the jobs API
and the attempt-scoped endpoint; it is not recoverable.
That leaves the row in U2 original condition --- a selector with no
fragments, unmatchable --- produced by exactly the mistake U3 is named
for. This is the fourth time this project has lost fragments this way,
and the first time I did it while holding the correction in my own
hands: I had corrected two other lanes for it earlier in the same
session.
Numbered U8, not U6, deliberately. U6 and U7 are reserved for the two
wall-clock rows on worker-identity-stage1, which renumbered into that
range when #229 took U4/U5. Taking U6 here would recreate the
duplicate-id collision that rebase already produced once, through the
same mechanism --- two lanes appending rows with no textual conflict.
The row is kept despite being unmatchable because of what it implies
together with U4 and U5: three distinct macOS selectors reddening in
one session points at a background failure rate on that platform rather
than three independent test bugs. That matters beyond bookkeeping,
because it undermines the equal-rate assumption behind any argument
about which branch a failure happened to land on --- including the one
currently being used to weigh #231.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The pre-rebase warning was right, and the mechanism is worth recording
because it is the quiet kind.
gate-protocol-build landed its own U4 and U5 in #229. On this rebase
git merged docs/ci-red-signatures.md WITHOUT A CONFLICT --- the two
lanes appended their rows in different places, so there was nothing
textual to resolve --- and produced two ### U4 and two ### U5 headings
describing entirely different incidents. No marker, no complaint.
That is the failure the matching rule exists to prevent, arriving
through the one path a careful conflict resolution would never catch:
there was no conflict to resolve.
Renumbered across all four sites the warning enumerated: both headings,
the prose relation-to-U4 field inside what is now U7, and the
active-work.md reference. Ids verified unique afterwards rather than
assumed.
The warning block itself is retired in place, replaced by a note saying
what was done and why, so the next reader sees a completed action
rather than an outstanding one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
gate-protocol-build independently defines its own U4 and U5 --- a macOS
lua54 PTY-resize failure and a Ctrl-C-as-SIGINT failure --- and it
merges first, so on main those ids are taken.
A rebase that resolves the textual conflict without renumbering leaves
two different incidents sharing an id, which is precisely the failure
this file matching rule exists to prevent. The registry authority rests
on ids meaning one thing.
The warning enumerates all four sites rather than saying "renumber the
rows", because one of them is a prose cross-reference inside U5
relation-to-U4 field and another is in active-work.md --- both easy to
miss when the conflict presenting itself is two adjacent headings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The header and the mechanism boundary were corrected last round; the
TABLE still numbered two. It called the 2026-08-09 worker run
"Occurrence 2", omitted the 2026-08-06 CRDT occurrence from the
enumeration entirely, and concluded "Two occurrences establish
intermittence" --- in the row I had just rewritten because it omitted
that same occurrence.
That is the head-and-body split this session keeps reproducing, this
time inside a single table, in the row whose whole purpose is to be the
authoritative account of what is known.
The row now enumerates all three, and says which one carries the most
weight: the 2026-08-06 CRDT run, because it shows the failure is not
confined to one feature flavor and can take the raw and canonical
selectors at once. That is a fact neither of the other two supplies.
Also corrected: "what is NOT: any mechanism, still" is now "no
mechanism is ESTABLISHED", because one IS proposed --- read-before-
write on the child output, the R4/R6 readiness family. Proposed is not
confirmed, and the row says so rather than flattening the distinction
in either direction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Two corrections, both mine, both the same failure the row exists to
warn about.
First, this is at least the THIRD occurrence, not the second, and the
fragment was not newly captured. docs/active-work.md records a
2026-08-06 loaded --features crdt run failing this selector AND
m6_1_pty_canonical_mode_keeps_kernel_echo with the same stty -a output
was: "" --- and it already proposed a mechanism family, read-before-
write on the child output, the shape of R4 and R6. So the row claim
that no mechanism had been proposed was false of the tree it was
written in. The evidence was in this repository the whole time; I wrote
a registry row without reading the registry neighbour.
Second, the fragment does not show what I said it showed. The test
inspects collect_stdout(&evs) after drain_until --- what the SUPERVISOR
collected. It cannot distinguish stty never writing from the PTY
dropping the bytes from event collection missing them. I wrote "stty
produced no output at all", which asserts a mechanism the test cannot
see, in the same row that says no mechanism is established.
What survives is narrower and still worth having: this is not a termios
failure, since nothing observed shows echo configured wrongly. Which of
child-never-wrote, delivery-lost, collection-missed is open.
The control changes accordingly. Sampling the collected string more
times cannot separate those three however often it fails; the next
occurrence needs the full process event stream and the child exit
disposition captured, cross-checked against the R4/R6 readiness family
that the 2026-08-06 entry already implicates.
The gate conclusion is unaffected: a markdown-only delta cannot cause a
PTY failure, so worker identity code is excluded as a cause.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The worker-identity tip gate went red on step 03-lib with one failure:
m6_1_pty_raw_mode_disables_kernel_echo. U2 already had that exact
selector but no fragments, so it could not be matched. Reading the
durable gate log rather than filtering a rerun supplies them.
The fragment reframes the failure. stty -a returned the EMPTY STRING,
not a wrong mode --- so this is not raw mode failing to disable echo,
it is stty producing no output at all, which points at PTY or spawn
readiness under load rather than termios handling. The assertion own
message is misleading on exactly that point, and anyone diagnosing it
from the message will look in the wrong place.
Occurrence 2 also EXCLUDES the change under test, which occurrence 1
could not. The tree carried zero code change since a 13/13 green run on
this same lane --- the only delta was three lines of markdown. A docs
edit cannot break a PTY test, so the diff is ruled out as a cause
rather than merely doubted. Isolated rerun passes in 0.01s.
Still no mechanism, and the row says so. Two occurrences establish
intermittence and a load correlation; neither establishes cause. The
row now names the discriminating control for a third: loop the selector
under synthetic load logging stty output every iteration, since whether
stty is empty EVERY time it fails is what separates a readiness race
from a termios one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The two reds recorded a moment ago were followed by a full green run of
the same command on the same tree — all 13 steps, log
`20260809T200907Z-2672209`. Both the lane entry and U5 now carry that,
because a signature row that records only the reds overstates them: the
green rerun is part of the evidence, not a reason to delete the row.
The row stays live and stays U-classified. Three load-sensitive render
budgets going red one per run and then green is consistent with a loaded
machine and with nothing else in hand; it is not a measurement of one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
The lane entry gains round 3: the wrong-surface diagnostic, why the job
and process refusals now say different things, the anti-collapse test
and its three mutation checks. Written here rather than left in the
commit message because this file is what a recovering agent reads.
`docs/ci-red-signatures.md` gains **U5**. Two consecutive `scripts/gate`
runs of the same command, on the same tree, red on step `12-sweep` with
a DIFFERENT wall-clock render-budget test each time — 224ms and 258ms
against a 200ms budget, 114ms against a 100ms budget, at load average
12.9/23.9 with sibling worktrees building. Each passes in an isolated
rerun of its own selector, and no selector reds twice.
The rotating selector is the signature, and it is a stronger one than
any single test name: a regression that moved between three unrelated
render paths on an unchanged tree is far less likely than one loaded
machine. The observing diff is two string literals, their doc comments
and one test, and touches no render path at all.
Kept separate from U4 rather than merged. U4 is two budget tests in
`04-lib-crdt` failing TOGETHER; this is three render-budget tests in
`12-sweep` failing ONE PER RUN. Merging them would assert a shared
mechanism nothing in hand shows, and the load confound stays unmeasured
in both — a rival explanation, not a finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Two files, no code.
## `docs/active-work.md` — review round 2
The lane's volatile block gains the round-2 record: the three findings,
why P2a's fix is a mapped diagnostic and P2b's is escaping at
presentation rather than rejection at the registry, and the seven new
mutation checks. The gate outcome is recorded with its step counts and
the two stop-signal facts — `journey_acceptance` 47/47 and all three
`#pmacs.process.list()` leak detectors byte-identical to `main`.
It also records what P2a's audit found and did NOT fix:
`pmacs.process.spawn`'s other string fields still convert generically.
That is pre-existing and out of this lane's diff, and it is named so it
is not silently inherited by whoever reads the fixed `purpose` read and
assumes the rest matches.
## `docs/ci-red-signatures.md` — R7's third occurrence, and U4
**R7 reproduced, and the control the second-occurrence note prescribed
finally discriminated — against its own hypothesis.**
Occurrence 2 left exactly one causal path open: the observing lane had
added a GPU-heavy `render_offscreen` test to the same binary, and
contention with a one-second socket handshake was plausible. That note
prescribed the control to run if a third occurrence landed — with the
added test removed, not at the merge base. A third occurrence landed, at
the gate's `gpu` step, with all three fragments verified against the
durable log.
The control was run. **Ten full `-p pmacs-gpu` runs with the added test:
10/10 green. Ten with it `#[ignore]`d, nothing else changed: 1 failure
in 10, all three fragments present.** Removing the suspect made the
failure more frequent, so the concurrent-test path is excluded — no
contention story from that test survives that direction.
The more useful result is the rate. This is the first rerun in R7's
history to reproduce anything at all, and it puts the failure at roughly
1-in-10 under ordinary `-p pmacs-gpu` load. Three sightings were not
enough to bisect a handshake; 1-in-10 is. The row now says so, and tells
the next agent to instrument which side closes the pipe rather than
re-run for green.
The lane is still not attributed — now for a measured reason rather than
an argument from diff shape: the arm without the lane's only
`pmacs-gpu` addition is the arm that went red.
**U4** records the other two reds from that same gate run:
`criterion_1_end_of_line_typing_completes_sub_frame_per_keystroke` and
`composition_overhead_under_ten_percent`, both wall-clock budget
assertions, failing together in `04-lib-crdt` and both green in
isolation and in the next full run. Fragments captured, so unlike U1–U3
it is matchable — it is a `U` row for want of a mechanism, not for want
of evidence. The signature named is **the pair**: two budget tests
failing in one run and neither in the next is far more likely to be one
loaded machine than two simultaneous regressions, and a future run that
reds only one of them is a different incident.
Neither row claims harmlessness, and the concurrent-worktree load
confound is recorded as a rival explanation rather than as a finding,
because it was not measured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
`attach::tests::managed_retry_survives_transients_and_uses_the_successful_stream`
failed once at this lane's `scripts/gate` **`gpu` step** on 2026-08-09.
Judged against this file rather than rerun-and-shrugged.
**It matches R7 on all three of its required fragments**, verified rather
than inferred:
transient sequence must attach: Attach(Handshake(Io(Os {
code: 32, kind: BrokenPipe, message: "Broken pipe" })))
**That capture is the point.** U2 and U3 both record the identical loss —
"output was filtered to the `FAILED` line" — and U3 says outright that
the recurring mistake was its author's, twice, with a mechanical fix:
read the durable log, never the live stream. The gate writes
`NN-gpu.log` for exactly this, and reading it turned what would have been
a third unjudgeable `U` note into a second occurrence of a row that had
one.
The flavor is a third one (`PMACS_REQUIRE_GPU=1 cargo test -p pmacs-gpu`,
neither occurrence 1's `--features crdt` sweep nor U3's default-features
workspace sweep). Recorded because this file's own R2 worked example
treats flavor as outside matching.
**The merge-base control R7 asked for was run, and it settles nothing.**
15 runs at `4bc55e8`, green — but the observing branch was green over 30
runs too (15 isolated selector, 15 full suite), so neither side
reproduced and the comparison separates nothing. Logged as a null result,
not as exculpation. Per the rerun rule, all 45 green runs establish
**intermittence only**.
**And one causal path is named rather than dismissed:** this lane adds a
GPU-heavy `render_offscreen` test to `pmacs-gpu`'s test module. It
touches no `attach.rs`, no protocol and no wire — but it does add a
concurrent test to the same binary, and the failing test is a socket
handshake on a one-second deadline. Contention is a plausible
`BrokenPipe` mechanism and 30 green runs do not exclude it. The row now
says what the discriminating control would be if there is a third
occurrence: remove the added test, not go to the merge base.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Three corrections. The first two are the same error in two rows.
I wrote each control as if it DECIDES, when each is informative in only
one direction.
U4: extending the deadline establishes "emitted late" IF the clear
arrives. If it does not arrive, that establishes only "not observed by
the longer deadline" --- not "never emitted" --- because transport loss
produces the same absence. No deadline, however long, separates
non-emission from transport loss. That needs producer-side emission
evidence, did pmacs write the clear, cross-checked against the
collected stream. The row now states both branches and names what the
negative branch cannot conclude.
U5: one isolated run cannot decide whether the gate suite is
implicated. A matching isolated RED proves the gate suite is not
necessary for the failure. An isolated GREEN proves nothing beyond that
run, because the failure is intermittent and absence under one run is
not evidence of dependence. I had written it as though either outcome
settled the question.
This is worth naming as a class rather than two typos: a control whose
positive branch is conclusive and whose negative branch is not, written
up as though both were, is how an inconclusive result gets recorded as
an exclusion. Two rows in this file had it.
Third, minor: "three docs" was accurate at the occurrence tip and is
not now --- #229 has since added this registry file. Replaced with
"documentation" so the claim does not rot again with the next commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai
Four review findings, all mine.
Normalization. The byte count and the :LINE suffix are
occurrence-specific --- the count is the collected suffix length, which
varies per run, and the line moves with the file --- so neither can be
a required fragment. Both rows now carry stable fragments and an
explicit NOT-fragments field naming what must not be matched on.
The LuaJIT-pass argument was the rejected overreach again. I used a
passing sibling leg as a STRUCTURAL exclusion; a deterministic defect
can be Lua-flavour-specific, so it is corroboration only. The row now
says so in those words. The real grounds are stronger anyway and were
sitting there: the workflow never invokes scripts/gate, and
full_grid_resync_acceptance runs BEFORE the changed gate suite, which
closes even the leaked-state path.
Three contradictions inside U4, each removed.
"It never emitted the blank" asserts a mechanism the next field
simultaneously calls open. Now: no blank was OBSERVED after the mark
within the deadline.
The 25,362 bytes were not a capped window. suffix.len() is the ENTIRE
post-mark output; only the displayed head is truncated, to 400 bytes.
Verified in the test source. So my control --- capture the full stream
rather than the window --- was solving a gap that does not exist. The
gap is arrival TIME, and the control now instruments that instead.
The ~20s failure duration IS the fixed Duration::from_secs(20) timeout,
so the ratio against a fast pass is mechanically determined and is not
independent timing evidence. Also verified in the source.
U5 control relabelled. Running m5_8_acceptance alone decides whether
the gate suite is implicated --- cross-suite attribution --- and
nothing more. It cannot separate injected-before-raw-mode from
raw-mode-lost from a third cause, and another isolated pass cannot
either however often it is repeated. Mechanism discrimination needs
readiness and raw-mode state observed AT INJECTION, which is now a
second, separately labelled control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bqGA6s9tTUFzYpbeW3tai