Commit Graph

5 Commits

Author SHA1 Message Date
Levi Neuwirth db24abb64e
feat(lsp): D3 --- the file watcher stops sleeping and walks once per scan (#233)
Implements docs/lsp-file-watch-d3-framing.md revision 4, approved
2026-08-11 with the four rulings adopted as proposed: the honest bar
(absent at idle, one attributable job per concurrently due group), no
exclusions by default, server root_uri -> cwd -> attachment fallback,
and constants rather than config keys.

pmacs.fs.walk_tree: the whole recursive tree as ONE cancellable job
(JobKind::FsWalkTree, reply reuses ReplyKind::ReadDir --- identical
payload shape, so the Lua boundary needs no second conversion). Names
are base-relative; symlinks recorded, never traversed; an unreadable
subdirectory skips its subtree (scan_tree's pcall behaviour); only the
root failing to open fails the walk; the cancel token is polled once
per directory. Eight Rust unit tests, including flat-directory entry
parity with read_dir_blocking and the two review-round cancellation
cases (empty-tree pre-cancel; mid-walk via the cfg(test) entry hook).

The watcher itself is rewritten as the framed group scheduler. No
sleeps anywhere: one process.after-tick subscription (installed once
and guarded --- pmacs.hook.remove does not exist) drives every
(server, base) group's deadline off monotonic_ms, autosave's Q#AS2
idiom. The old design held one pool thread per sleeping watcher and
allocated 1 sleep + D read_dir jobs per watcher per tick --- 1,326
per tick for rust-analyzer's six watchers on this 220-directory
checkout. At idle there is now NO running job, which is also the
strongest witness in the suite: activity_summary settles to None, and
that assertion is unwritable under the old design.

The scheduler is the framing's state machine, all three review rounds
included: single-flight per group with generation-checked completions;
deadlines advanced from completion; the round-3 three-arm completion
partition (success / stale-or-retired / live non-success, with the
failure latch and quiet cancellation); joins wake the group, queue
exactly one follow-up mid-walk, and never reset the backoff curve;
per-watcher baselines --- the first snapshot whose WALK STARTED after
the join; membership captured at scan start; per-member cancellation
recheck at emit through the preserved _after_scan_for_tests seam;
backoff 250ms x2 to a 4s cap, reset by any emitted change; retirement
cancels the in-flight walk cooperatively.

Verification: eighteen acceptance tests. The six #234 tests are
byte-unchanged and green. Ten witnesses cover the framing's plan (the
review rounds added the fallback-determinism and root-boundary pair,
making twelve):
idle absence (and never a sleep purpose), one walk job per scan on a
twelve-directory fixture, join-wakes plus the registration epoch,
queued baseline for a mid-walk join (driven by saturating the worker
pool so the walk genuinely queues), single-flight under a withheld
completion pump, retirement and rebaseline through the fake's
unregister/re-register triggers, live cancel via pmacs._async._cancel
on the queued job, live failure with the once-per-error latch and the
preserved-snapshot recovery (DELETED for the pre-failure file is only
derivable from the retained snapshot), backoff shape from seam
timestamps, and the configured-root base.

Every witness was mutation-tested. Two findings from the bites:

- Retirement is DOUBLE-ENFORCED (unregister path and post-scan sweep)
  and biting either copy alone is masked by the other; only biting
  both goes red. Kept deliberately: the sweep covers seam-cancelled
  members, the unregister path covers idle groups whose next deadline
  is seconds away.
- The first idle probe was VACUOUS: it read pmacs.async instead of
  pmacs._async, errored, and the unwrap_or_default made every sample
  read as "absent". The probe now expects rather than defaults, so a
  broken probe is a red test, not a green lie.

One environmental fact, recorded in the lane: an empty stray /tmp/.git
(since removed) made project detection root every markerless tempdir
fixture at /tmp, which under Q#D3-3 the watcher then faithfully
watched. A markerless-fixture red that looks like a watcher bug may be
an ancestor marker.

A pre-commit review round found four blockers, all fixed here:

- walk_tree checked cancellation only inside its entry loops, which an
  EMPTY tree never enters --- a pre-cancelled queued walk returned an
  empty SUCCESS, which the success arm would commit and diff into a
  deletion storm. Cancellation is now checked before opening and
  before returning, cancellation outranks a missing-root error, and a
  unit test pins both.
- The neither-root-nor-cwd attachment fallback was still pairs-order
  nondeterministic --- the exact accident D3 set out to remove, behind
  a comment claiming otherwise. It now takes the lexicographically
  smallest attachment directory. Verified at the spawn sites: every
  server spawned with an attached file gets cwd = root, so the arm is
  defensive and unreachable through production spawning --- which is
  also why it carries no through-the-server witness.
- A base at the filesystem root joined as //path (and file:////path in
  URIs). Both join sites now go through join_under, the root-aware
  idiom dired's handler already uses, and the dir-of capture for a
  root-level file ("" from the match) normalizes to "/".
- The walk-count and scan-times probes defaulted on error, so two
  broken probes could compare equal and pass the retirement witness.
  Every probe now expects --- a broken probe is a red test, the same
  correction the vacuous idle probe forced.

A second pre-commit round found three more, all fixed here:

- Mid-walk cancellation was UNWITNESSED: both Rust cancel tests
  pre-cancelled and the acceptance test cancelled a queued walk, so
  deleting the internal polls left every test green. A cfg(test)
  entry hook now flips the token at an exact entry boundary and the
  witness asserts the walk stopped NEAR it (bound on entries
  processed), which is what discriminates the polls from the
  entry/exit checks. The retirement witness now holds a walk in
  flight across the unregister and asserts the job settles cancelled.
- The "unreachable fallback" claim was WRONG: pmacs.lsp.spawn may
  omit both cwd and root_uri, and ensure_server adopts such a live
  server for markerless files (root_uri and key_uri both nil). The
  lexicographic-minimum fallback now has a through-the-server
  witness: five sibling directories, the minimum opened last ---
  five, because with two the build's hash order coincided with the
  lexicographic answer and the first-pairs bite survived.
- The root-boundary joins gained a witness through exported
  production functions (the _deliver_status pattern): the matcher and
  URI builder driven at base "/", where reverting either join_under
  call makes the anchored glob refuse //hit and the URI grow a fourth
  slash. No fixture can walk / for real.

Verification totals after both rounds: eight walk_tree unit tests,
eighteen acceptance tests (six byte-unchanged, twelve witnesses), all
mutation-verified.

No wire change, no PROTOCOL_VERSION bump; walk_tree is an fs binding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:59:28 +02:00
Levi Neuwirth ec3473d598
docs: close D3 scheduler non-success transition
Revision 4 absorbs review round 3. A live walk cancellation or failure
now has an explicit group-state transition: no snapshot, epoch, or emit;
the prior state and backoff survive; in-flight clears; queued joins run
immediately; otherwise the group reschedules. Distinct failures report
once until a successful scan clears the latch.

Correct the queued-baseline latency bound and add live-cancel and
live-failure witnesses. Record the round in the active-work lane; the
four user rulings still block implementation.
2026-08-11 13:00:41 +02:00
Levi Neuwirth 9c644b0ae7
docs: D3 framing revision 3 --- the scheduler becomes a state machine
Review round 2: two P1 design gaps and one P2 overclaim, all in the
cadence revision 2 introduced.

A join now wakes the group: next_scan_at pulls to now, an in-flight
walk queues exactly one immediate follow-up, and a baseline is only a
snapshot whose WALK STARTED after the join --- an in-flight walk may
have passed a directory before a pre-join file appeared there, so its
snapshot as a baseline would turn that file into a false CREATED.
Without the wake, a backed-off group folds post-registration files
into the baseline and never reports them; today registration scans
immediately and the coalesced design must not regress that.

The group is now a defined state machine: single-flight per group
(overlap unrepresentable, not avoided), deadlines advanced from
completion (a walk outliving its interval degrades to back-to-back
scans, never overlap), stale completions rejected by generation
(#234 P2 at group scope), and retirement --- last member gone or
server death --- that cooperatively cancels the walk. Cancellation
therefore enters walk_tree contract and tests; polling the cancel
token between directory reads is the established job shape. The
after-tick subscription installs once and guards, because
pmacs.hook.remove does not exist (the P3 gap).

Q#D3-1 restated honestly: groups key on (server, base), so several
can be due on one frame and per-group single-flight still permits N
jobs. The bar offered is absence at idle plus one attributable job
per concurrently due group, with a global scan queue as the
alternative if one-at-a-time must be guaranteed.

Six round-2 witnesses join the plan: join-wakes, no-overlap,
retirement, queued-baseline, plus the round-1 epoch and idle pair.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 12:48:41 +02:00
Levi Neuwirth 0aac3b8992
docs: D3 framing revision 2 --- round 1 removed the sleep, the guess, and the skip
Review round 1 (five findings, four P1) and revision 1 did not survive
it. The findings are recorded in the document; three are cases where
the framing reasoned from the wrong mechanism.

The promised idle state was impossible: pmacs.workers.sleep is a
RUNNING job for its whole duration --- dispatched onto the worker
pool, sleeping in 1ms slices on one of available_parallelism - 1
threads --- and activity_summary counts every running job. A 4s
backoff sleep renders as a constant "sleep 4000ms", and filtering it
would touch the instrument. Revision 2 removes the sleep from the
design: one process.after-tick subscription owns every group's
schedule via monotonic_ms --- autosave's Q#AS2 idiom, whose own
comment names the pool-thread hazard. Waiting now allocates no job
and no thread; the indicator is absent at idle by its None-at-zero
contract, which also becomes the strongest witness in the plan.

The scan root is the server's, not a guess: root_uri then cwd then
the attached-file fallback. project.detect was wrong by the tree's
own testimony --- server rooting honors configured strings and
resolvers first, and texlab's Q#LX2 documents a root that detect can
never produce.

Coalescing gained delivery semantics: shared snapshots, per-watcher
baselines (the first snapshot completed after join), membership
captured at scan start, cancellation rechecked per watcher at emit
(#234's P2 rule, per member). Two epoch witnesses join the plan; the
six existing tests do not cover this and were never claimed to.

Exclusions default to NONE. Any unconditional skip deviates from the
registered glob contract (**/-leading globs can match under .git/,
and a server may register .git/HEAD outright), and walk_tree removes
the job-count economics that made skipping look necessary: 80% of
jobs becomes readdir syscalls inside one job.

Arithmetic corrected to this checkout: D = 220, so 221 jobs per
watcher per tick and 1,326 for rust-analyzer's six --- six of them
pool-thread-holding sleeps. Revised steady state: zero jobs at idle,
one walk_tree job while a scan runs.

Q#D3-1..4 rewritten accordingly; all four still block implementation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 12:37:40 +02:00
Levi Neuwirth bfa1e77d1b
docs: frame D3 --- the file-watch polling cost (issue #233)
Framing and lane update at the branch's first commit, portable during
review. Revision 1 is a DRAFT and no implementation may start from it.

The remainder of issue #233 after #234: the watcher is correct and
still walks everything. The framing re-verifies the cost model at
add0ba1 and adds two facts the issue does not carry:

- On this checkout .git alone is 177 of 220 directories --- over 80%
  of every walk --- and the machine does not even have an in-tree
  target/ (external CARGO_TARGET_DIR). The issue's 187/46 numbers were
  measured with one.
- The string-form base is NONDETERMINISTIC: resolve_watcher takes the
  directory of whichever attachment pairs() yields first, so which
  tree a bare-string watcher can see depends on table hash order. #234
  made matching correct per base; which base is still accidental.

Design space A-E with the trade-offs stated: coalesce per (server,
base); a Rust walk_tree primitive (one job per scan instead of one per
directory --- no new crate, no wire change); an ignore list, with the
target/ staleness hazard named (rust-analyzer's **/*.rs covers
OUT_DIR outputs, so an aggressive default trades churn for staleness
in exactly the server the issue is about); idle backoff; and kernel
notification, deliberately staged separately because A-D are pure wins
it does not obsolete and a new-crate decision deserves its own
framing.

Proposed Stage 1: walk_tree + coalescing + backoff, VCS-only skip
default, project-root base preference. At rest on this repo: from
~1,100 jobs per scan-bound tick to one job every four seconds, named
for its root.

Four rulings block implementation (Q#D3-1..4): the acceptance bar
(one attributable blip at idle, not silence --- silence is Stage 2),
the skip-list default and where it lives (a list-valued setting is
not expressible in today's ConfigValue), the string-form base change,
and knobs vs constants.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 10:03:57 +02:00