test(process): read the pid without ticking in the fast-exit tests

The parallel workspace sweep failed
observing_the_leader_does_not_consume_the_exit_event with "process
ProcessId(26) is not running". A real defect in the test, not a flake.

The helper that fetched the pid drained for the Started event, and
draining ticks. A tick can observe an immediately-exiting child and
transition the record out of Running, after which signal returns "is not
running" and never reaches the diagnostic -- so the loop spun to its
10 s bound and panicked. It passed standalone because the drain returned
on Started before poll_one saw the exit; only the sweep's load shifted
the timing enough to lose that race.

Fast-exiting children now read the pid straight from the supervisor
record, which does not tick. The bounded loop also fails fast when the
record has left Running, so a future recurrence is diagnosed in one line
rather than surfacing as a timeout.

Verified under matched load: 15/15 green with all 16 cores saturated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HZjWMjwPXhPbt9upku9mCk
This commit is contained in:
Levi Neuwirth 2026-07-25 22:25:05 -04:00
parent 18d481b046
commit 00cc615db5
1 changed files with 32 additions and 3 deletions

View File

@ -2296,7 +2296,27 @@ mod tests {
(id, spawn_started_pid(sup, id))
}
/// Drain until `Started` and return the OS pid it carries.
/// The OS pid straight from the supervisor's own record, WITHOUT
/// ticking.
///
/// `drain_until` ticks, and a tick can observe a fast child's exit
/// and transition the record out of `Running` — after which
/// `signal` returns "is not running" and never reaches the
/// diagnostic at all. Any test whose child exits promptly must read
/// the pid this way. (Found by the parallel workspace sweep: the
/// drain-based helper raced only under load.)
fn record_pid(sup: &ProcessSupervisor, id: ProcessId) -> u32 {
match sup.processes.get(&id).expect("record").state {
ProcessState::Running { pid, .. } | ProcessState::Exiting { pid, .. } => pid,
ProcessState::Starting => panic!("spawn has not reported a pid yet"),
ProcessState::Terminated(_) => {
panic!("the record already left Running; the pid is unavailable")
}
}
}
/// Drain until `Started` and return the OS pid it carries. Safe
/// only for children that outlive the drain; see [`record_pid`].
fn spawn_started_pid(sup: &mut ProcessSupervisor, id: ProcessId) -> u32 {
let evs = drain_until(sup, id, Duration::from_secs(5), |evs| {
evs.iter()
@ -2331,6 +2351,11 @@ mod tests {
if err.contains("leader=exited(") {
return err;
}
assert!(
!err.contains("is not running"),
"the record left Running before the diagnostic could run, so \
this test never exercised it: {err}"
);
assert!(
Instant::now() < deadline,
"leader never observed as exited within {timeout:?}: {err}"
@ -2430,7 +2455,9 @@ mod tests {
let mut spec = ProcessSpec::new("diag-exited", "/bin/sh");
spec.args = vec!["-c".into(), "exit 3".into()];
let id = sup.spawn(spec).expect("spawn");
let pid = spawn_started_pid(&mut sup, id);
// NOT `spawn_started_pid`: draining ticks, and this child exits
// immediately.
let pid = record_pid(&sup, id);
let err = terminate_until_leader_exited(&mut sup, id, Duration::from_secs(10));
@ -2506,7 +2533,9 @@ mod tests {
mode: TerminalMode::Canonical,
};
let id = sup.spawn(spec).expect("spawn");
let _ = spawn_started_pid(&mut sup, id);
// NOT `spawn_started_pid`: draining ticks, and a tick can reap
// this immediately-exiting child before the diagnostic runs.
let _ = record_pid(&sup, id);
// Drives `observe_leader`, which try_waits the real PTY child
// for the first time and reaps it.