docs(ci-reds): U9's control is VOID --- cargo runs test binaries serially

U9's "structural difference worth testing next" claimed that
`cargo test --workspace` runs many test binaries concurrently while
`--lib` runs one, and derived its discriminating control from that:
"pin test-binary concurrency to 1". The premise is false. Cargo runs
test TARGETS serially, one executable at a time, so that concurrency is
already 1 and the control pins nothing.

Measured in this project's own gate logs rather than asserted from the
cargo book: `20260831T093655Z-857818/07-sweep.log` alternates `Running`
and `test result:` strictly --- 119 to 121 markers, ZERO cases of one
binary starting before the previous reported. The pattern is `RTRTRT`.

That falsifies a premise two rows rested on, so U12's family paragraph
is corrected too. What survives is smaller and still true: a sweep is a
long sequence of binaries, so a budget inside it runs at an arbitrary
point in a multi-minute step. The family still should not consume review
rounds --- but it now needs a control someone has to design.

U17 no longer claims `--test-threads=1` exercises U9's control. It is a
different knob: it serializes test FUNCTIONS within one executable. Its
candidate mechanism is narrowed to match --- removing sibling test
functions removes ONE source of contention, which supports neither
"fastest" nor "narrowest".

R6's block drops two overclaims: a PR run CAN show the identical red
(only the main dispatch establishes it on the merge base), and this was
not the dispatch key's first use --- #245's D2/D3 dispatched three runs
right after it merged. It is the first use for a live merge-base
control.
This commit is contained in:
Levi Neuwirth 2026-08-31 15:05:36 +02:00
parent 088f24e1bb
commit a7c4b3adec
No known key found for this signature in database
2 changed files with 74 additions and 44 deletions

View File

@ -363,9 +363,8 @@ The lint itself is `#[allow]`ed with a reason — collapsing the three
hide that `(forward, empty, None)` and `(history, empty, Some)` are hide that `(forward, empty, None)` and `(history, empty, Some)` are
valid for opposite reasons. **The gap is not fixed here**: adding a valid for opposite reasons. **The gap is not fixed here**: adding a
second clippy flavor to `scripts/gate` is a change to shared second clippy flavor to `scripts/gate` is a change to shared
infrastructure and belongs in its own lane, alongside U9's still-unrun infrastructure and belongs in its own lane. Recorded so the next lane
discriminating control. Recorded so the next lane touching touching crdt-gated code does not rediscover it at CI.
crdt-gated code does not rediscover it at CI.
**Four registry rows moved on this lane:** **Four registry rows moved on this lane:**
@ -415,9 +414,10 @@ crdt-gated code does not rediscover it at CI.
**The log was read before anything was rerun**, which is U3's lesson **The log was read before anything was rerun**, which is U3's lesson
and U8's fourth-violation warning finally honoured on a macOS job. A and U8's fourth-violation warning finally honoured on a macOS job. A
**merge-base control was dispatched** at `aae5b35` rather than **merge-base control was dispatched** at `aae5b35` rather than
arguing from an unrelated diff — **the first real use of the arguing from an unrelated diff. **Not the `workflow_dispatch` key's
`workflow_dispatch` key #245 landed**, and exactly the case U11 first use** — #245's own D2/D3 witnesses dispatched three runs right
motivated it for. It came back **green on the macOS legs**, so the after it merged — but **the first use for a live merge-base
control**, which is the case U11 motivated it for. It came back **green on the macOS legs**, so the
inference it could have supplied is **unavailable**; recorded as a inference it could have supplied is **unavailable**; recorded as a
null result, as R1's row had to record its own; null result, as R1's row had to record its own;
- **U17, new** — that same control run **redded `Test (crdt)` on `main` - **U17, new** — that same control run **redded `Test (crdt)` on `main`
@ -426,10 +426,11 @@ crdt-gated code does not rediscover it at CI.
way to R1 and R5 — not a missed deadline. What `got ok` proves is way to R1 and R5 — not a missed deadline. What `got ok` proves is
narrow: the predecessor **completed successfully before cancellation narrow: the predecessor **completed successfully before cancellation
took effect**, which does not say when the supersede arrived. The job took effect**, which does not say when the supersede arrived. The job
runs `--test-threads=1`, the condition **U9's still-unrun control runs `--test-threads=1`, which serializes test **functions within one
names**. A PR run could show this failure too; what only a `main`-side executable** — **not** the test-**binary** concurrency U9's control
run establishes is that it fails **on `main`**, with no observing named, and cargo runs binaries serially anyway. A PR run could show
branch to suspect. this failure too; what only a `main`-side run establishes is that it
fails **on `main`**, with no observing branch to suspect.
U14 and U15 are two rows rather than one because the second run's U14 and U15 are two rows rather than one because the second run's
selector set had **rotated**, and this registry matches on the exact selector set had **rotated**, and this registry matches on the exact
@ -630,10 +631,14 @@ from #171 and #215.
the full 36-test binary both passed immediately afterwards — the full 36-test binary both passed immediately afterwards —
intermittence only. This lane changes neither the gate script nor intermittence only. This lane changes neither the gate script nor
that acceptance binary; diagnostic hardening is a separate lane. that acceptance binary; diagnostic hardening is a separate lane.
- **Still owed, separately:** `workflow_dispatch` on `ci.yml`, and U9's - **Still owed, separately:** `workflow_dispatch` on `ci.yml` (**landed
discriminating control — pin test-binary concurrency to 1, then load as #245**), and U9's discriminating control — named since 2026-08-09,
a lone `--lib` binary — which has been named since 2026-08-09 and never run, and now **VOID**: its premise that `cargo test --workspace`
never run. runs many test binaries at once is false, cargo runs test targets
serially, so "pin test-binary concurrency to 1" pins something already
1. See the correction on U9 in `docs/ci-red-signatures.md`. **A
replacement control has to be designed**; the budget family no longer
has one written down.
## Panel-pointer replay (parent acceptance 48) — MERGED as #243 (`6c9bae6`) ## Panel-pointer replay (parent acceptance 48) — MERGED as #243 (`6c9bae6`)

View File

@ -269,9 +269,15 @@ nothing in the panel, process or terminal paths**. But "my diff looks
unrelated" is not evidence, so unrelated" is not evidence, so
[run 33375945966](https://github.com/levineuwirth/pmacs/actions/runs/33375945966) [run 33375945966](https://github.com/levineuwirth/pmacs/actions/runs/33375945966)
was dispatched at `aae5b35`, **the branch's exact merge base**, via the was dispatched at `aae5b35`, **the branch's exact merge base**, via the
`workflow_dispatch` key #245 landed for precisely this. **This is that `workflow_dispatch` key #245 landed for precisely this.
key's first real use**, and U11 — the row that motivated it — is why it
exists. **It is NOT that key's first use, and an earlier version of this block
said so.** #245's own owed witnesses D2 and D3 dispatched three runs
(`33307137965`, `33308891808`, `33308921103`) immediately after it
merged, and the first of those already found a red on `main`. What this
is: **the first use for a live merge-base CONTROL** — a contemporaneous
`main`-side run obtained to answer a specific branch-side red, which is
the case U11 motivated the key for.
**THE CONTROL LANDED GREEN on the macOS legs**, and the meaning was **THE CONTROL LANDED GREEN on the macOS legs**, and the meaning was
pre-registered above before the result was seen: pre-registered above before the result was seen:
@ -283,9 +289,10 @@ exactly what R1's row had to record about its own green control, and it
is recorded the same way here: **a null result, not an exculpation.** is recorded the same way here: **a null result, not an exculpation.**
**The control run was not otherwise clean, and that is its own finding.** **The control run was not otherwise clean, and that is its own finding.**
`Test (crdt)` **failed on `main` at `aae5b35`** — see **U17**. A red on `Test (crdt)` **failed on `main` at `aae5b35`** — see **U17**. **A PR
the merge base is not something a PR run can show you; it took a run can show the identical red**, and an earlier version of this block
dispatch on `main` to see it at all. denied it; what only the `main` dispatch establishes is that the failure
occurred **on the merge base**, with no observing branch to suspect.
**Circumstantial alignment with U8, deliberately NOT a merge.** U8 has **Circumstantial alignment with U8, deliberately NOT a merge.** U8 has
the same selector, panicking at the **same line** `:2454` with the the same selector, panicking at the **same line** `:2454` with the
@ -1179,7 +1186,7 @@ claim is the one a later reader would otherwise reach for.*
| **status** | **one occurrence; INTERMITTENT — the identical sweep command on the same tree was green (118 targets, 1928 passed, exit 0)** | | **status** | **one occurrence; INTERMITTENT — the identical sweep command on the same tree was green (118 targets, 1928 passed, exit 0)** |
| **what IS established** | intermittence, with the strongest available exclusion of the tree: green in two earlier steps of the **same run**, green isolated afterwards (`2 passed`, 1.70 s), green on a full sweep rerun. Both assertions are **timing-sensitive by construction** — one reads collected child output within a deadline, the other measures wall-clock composition overhead (observed 1.613× against a 1.10× budget; 61.3% dispatch and 124.6% realistic overhead) | | **what IS established** | intermittence, with the strongest available exclusion of the tree: green in two earlier steps of the **same run**, green isolated afterwards (`2 passed`, 1.70 s), green on a full sweep rerun. Both assertions are **timing-sensitive by construction** — one reads collected child output within a deadline, the other measures wall-clock composition overhead (observed 1.613× against a 1.10× budget; 61.3% dispatch and 124.6% realistic overhead) |
| **what is NOT** | cause, and the load confound is **partially measured but NOT controlled**. The failing sweep ran inside a full gate; the green rerun started at load average 1.98 with the 5-minute figure still at 8.03 from that gate. Different conditions is not a measurement of the mechanism, and this row does not treat it as one | | **what is NOT** | cause, and the load confound is **partially measured but NOT controlled**. The failing sweep ran inside a full gate; the green rerun started at load average 1.98 with the 5-minute figure still at 8.03 from that gate. Different conditions is not a measurement of the mechanism, and this row does not treat it as one |
| **the structural difference worth testing next** | `cargo test --workspace` runs **many test binaries concurrently**; `--lib` runs **one**. That is a difference in kind between the passing steps and the failing one, not merely a difference in load average — and it is the first candidate this family has had that is checkable rather than atmospheric. **Discriminating control:** rerun the sweep with test-binary concurrency pinned to 1, and separately run the `--lib` binary alone under synthetic load. A red under synthetic load at low sweep concurrency implicates load; a red at high concurrency and low load implicates the concurrency itself | | **the structural difference worth testing next — PREMISE FALSIFIED 2026-08-31** | This cell claimed `cargo test --workspace` runs **many test binaries concurrently** while `--lib` runs one, and derived a control from it: "pin test-binary concurrency to 1". **Cargo runs test targets SERIALLY**, one executable at a time, so that concurrency is already 1 and the control pins nothing. Measured in this project's own logs: `20260831T093655Z-857818/07-sweep.log` alternates `Running` and `test result:` strictly, 119 to 121, with **zero** overlapping starts. `--test-threads=1` is a *different* knob — it serializes test functions **within** one executable — and does not stand in for the control either. **The real difference between the steps is which binaries run and how long the whole step takes, not how many run at once.** A replacement control has to be designed; this row no longer has one |
| **relation to U2 — a NEAR MISS, do not match it there** | the PTY fragment is U2's exact family (`stty -a output was: ""`), but U2's selector field names only `m6_1_pty_raw_mode_disables_kernel_echo`. U2's occurrence 2 saw raw **and** canonical fail together; here **canonical redded alone and raw passed**, which U2's evidence has never shown. It is recorded here rather than folded into U2 so that the "canonical alone" case stays visible | | **relation to U2 — a NEAR MISS, do not match it there** | the PTY fragment is U2's exact family (`stty -a output was: ""`), but U2's selector field names only `m6_1_pty_raw_mode_disables_kernel_echo`. U2's occurrence 2 saw raw **and** canonical fail together; here **canonical redded alone and raw passed**, which U2's evidence has never shown. It is recorded here rather than folded into U2 so that the "canonical alone" case stays visible |
| **relation to U6 — its own instruction, honoured** | `composition_overhead_under_ten_percent` is one of U6's two selectors, and U6 says plainly: "If a future run reds **one** of these without the other, that is a different incident and should be judged as one." It redded without `criterion_1_end_of_line_typing…`, in a different step, at a far larger margin (1.613× here against U6's 1.297×). Judged as a different incident, as instructed | | **relation to U6 — its own instruction, honoured** | `composition_overhead_under_ten_percent` is one of U6's two selectors, and U6 says plainly: "If a future run reds **one** of these without the other, that is a different incident and should be judged as one." It redded without `criterion_1_end_of_line_typing…`, in a different step, at a far larger margin (1.613× here against U6's 1.297×). Judged as a different incident, as instructed |
| **what this row does NOT assert** | that the two selectors share a mechanism. They failed together once; they belong to different subsystems; and U7 already refused this exact merge for U6. The **co-failure inside one step with an in-run green control** is the signature — not either name, and not a shared cause | | **what this row does NOT assert** | that the two selectors share a mechanism. They failed together once; they belong to different subsystems; and U7 already refused this exact merge for U6. The **co-failure inside one step with an in-run green control** is the signature — not either name, and not a shared cause |
@ -1208,15 +1215,27 @@ green in the other run**.
| **relation to U6** | run B's selector is one of U6's two, redding **without** `composition_overhead_under_ten_percent`. U6 instructs that one-without-the-other is a different incident; honoured here | | **relation to U6** | run B's selector is one of U6's two, redding **without** `composition_overhead_under_ten_percent`. U6 instructs that one-without-the-other is a different incident; honoured here |
| **what this row does NOT assert** | a shared mechanism between the two rows, or any mechanism at all. **The signature is the rotation across an identical commit** — not either name | | **what this row does NOT assert** | a shared mechanism between the two rows, or any mechanism at all. **The signature is the rotation across an identical commit** — not either name |
**Why this family keeps recurring, stated plainly.** Every row in it is **Why this family keeps recurring — with its stated premise CORRECTED,
a wall-clock budget asserted **inside a workspace-wide parallel test because it was false.** Every row in it is a wall-clock budget asserted
run**. `cargo test --workspace` starts many test binaries at once, so inside a workspace-wide test run. This paragraph used to add that
each budget competes with the rest of the sweep in **every** run, "`cargo test --workspace` starts many test binaries at once". **It does
including the ones that pass. A 4.5% overshoot on a 1ms budget is not a not. Cargo runs test targets SERIALLY, one executable at a time**, and
signal about the code. **U9 already named the discriminating control** this project's own gate logs measure it: in
— pin test-binary concurrency to 1 and separately load a lone `--lib` `20260831T093655Z-857818/07-sweep.log`, 119 `Running` markers and 121
binary — and it remains unrun. Until it runs, this family should not `test result:` lines alternate strictly — **zero** cases of one binary
consume another review round. starting before the previous one reported. The pattern is `RTRTRT…`.
So the budgets do **not** compete with the rest of the sweep in the way
this family assumed. What is still true is smaller: a sweep is a long
sequence of binaries, so any budget inside it runs at an arbitrary point
in a multi-minute step, on whatever the machine is doing then. A 4.5%
overshoot on a 1ms budget remains not a signal about the code.
**And U9's named control does not discriminate what it claimed** —
"pin test-binary concurrency to 1" pins something that is *already* 1.
See the correction on U9 itself. This family still should not consume
another review round, but it now needs a control someone has to design,
not one already written down.
**Widening a budget is not the fix**, and R1 already rejected it. **Widening a budget is not the fix**, and R1 already rejected it.
@ -1519,21 +1538,27 @@ assumes a late arrival the assertion cannot see. A supersede that
arrived in time and whose cancellation simply did not take effect first arrived in time and whose cancellation simply did not take effect first
produces the identical message. produces the identical message.
**Candidate mechanism, stated as one.** The job runs **Candidate mechanism, stated as one — and stated smaller than an
**`--test-threads=1`**. A test that dispatches a job and then supersedes earlier version had it.** The job runs **`--test-threads=1`**, which
it "in flight" depends on the predecessor still being in flight; with no serializes the **test functions inside one libtest executable**. A test
other test competing for the runtime, the predecessor is at its that supersedes a job "in flight" depends on the predecessor still being
*fastest*, and the window in which it can be superseded is at its in flight, and removing sibling test functions from the same process
narrowest. That is a reading of the assertion and the flag together — it **removes one source of contention** for it.
is **not** a diagnosis, and nothing here rules out a real supersede
defect.
**Worth noting for U9.** U9's still-unrun discriminating control is That is all it supports. The earlier wording said the predecessor is at
"pin test-binary concurrency to 1". This job **already does that**, in its *fastest* and the window at its *narrowest*; neither follows. Other
CI, on every run. That does not run U9's control — different binary, contention remains — the rest of the machine, the CI runner's own load,
different selectors — but it does mean the single-threaded condition is and every other process — and nothing here measured the predecessor's
not hypothetical in this project, and a row now exists where it may be duration with the flag on versus off. It is **not** a diagnosis, and
load-bearing in the opposite direction from every budget row here. nothing rules out a real supersede defect.
**Worth noting for U9 — and NOT as an instance of its control.** An
earlier version said this job "already does" what U9's control asks. It
does not, and the distinction is the whole point of U9's premise:
**`--test-threads=1` serializes test FUNCTIONS within one executable; it
does not pin test-BINARY concurrency.** Those are different knobs. See
the correction recorded against U9 and U12 below, which is larger than
this note.
**Not R1 and not R5**, on this file's own matching rule: different **Not R1 and not R5**, on this file's own matching rule: different
selector, different module, different assertion. R5's row draws exactly selector, different module, different assertion. R5's row draws exactly