diff --git a/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md b/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md deleted file mode 100644 index 2756a638..00000000 --- a/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md +++ /dev/null @@ -1,43 +0,0 @@ ---- -id: itd-2609211335097114 -slug: a-stuck-model-never-holds-the-server-hostage-alice-asks-her -spec_id: null -kind: null -suggested_kind: null -reclassification_history: [] -builds_on: [] -severity: minor -impact: additive -origin: researcher-authored -production_mode: hand-written ---- - -# A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed. - -## Press Release - -> _Seeded from a quoted-text intent capture. Expand into the full press-release narrative before planning._ - -## Why This Matters - -A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed. - -## Mechanism - -> _Prompted (the claim-recording gradient): why the authors expect this to work, as a falsifiable "we expect X because Y" — not the outcome restated. Replace this line with the claim, or with the exact token `None stated.` alone on its line to record the claim as considered and declined._ - -## Scope Conditions - -> _Required (the claim-recording gradient): the population, platform, scale, or assumptions this claim holds under, one per top-level bullet — `abcd intent plan` stamps each with a persistent identity. Replace this line with those bullets, or with the exact token `None stated.` alone on its line._ - -## Acceptance Criteria - -> _Required (the itd-1 discipline): add at least one Given-When-Then bullet describing the verifiable bar for "shipped" before this draft can be planned._ - -## Open Questions - -_None recorded yet._ - -## Audit Notes - -_Empty. Populated by intent-auditor when intent moves to shipped/._ diff --git a/.abcd/development/intents/planned/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md b/.abcd/development/intents/planned/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md new file mode 100644 index 00000000..29514c07 --- /dev/null +++ b/.abcd/development/intents/planned/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md @@ -0,0 +1,164 @@ +--- +id: itd-2609211335097114 +slug: a-stuck-model-never-holds-the-server-hostage-alice-asks-her +spec_id: spc-2609211753023984 +kind: standalone +suggested_kind: null +reclassification_history: [] +builds_on: [itd-2609091301112705, itd-2609100457007827, itd-2609061441241254, itd-2609061441285238, itd-2609091412177263] +severity: minor +impact: additive +origin: researcher-authored +production_mode: hand-written +--- + +# A stuck model never holds the server hostage: a real request pre-empts idle work + +## Press Release + +A stuck model never holds the server hostage. + +Bob asks a question from his laptop on the network while the Dessau Server in +Alice's study is measuring a model in the background — loading it, even. He +gets his answer. The server notices that his request needs the memory its own +idle work is holding, stops that work at once, and loads the model he asked +for; the pause is a few seconds at most, and if the memory does not come back +within the server's own bound he is told, in the refusal, what held it and that +it was the server's work rather than the size of what he asked for. The +measurement the server abandoned loses nothing it had verified: it resumes +later from where it stood. What Bob never sees is a 503 he has to retry, and +what Alice never sees is a server that answers nobody for hours because it is +busy with a job it started itself. + +Alice can see the same thing from the control panel: a model held by idle +work says which job holds it and for how long, and one click frees it, as it +does today. A model she has pinned is the one thing idle work cannot give up +on her behalf: a request that needs a pinned model's memory is refused, +naming the holder, and the pin stays. + +## Why This Matters + +On 2026-09-21 the live server answered 503 to every chat request for six hours +because its context probe had queued a model the runtime cannot load and +retried it thirty-two times, each attempt holding the whole memory budget +(iss-2609211334563318, iss-2609211334570516, iss-2609211334576018; fixed in +0.9.3 by PR 144, which makes such a load fail in seconds, keeps a failed model +out of idle work, restricts the probe to chat models, and names the holder in +the refusal). What 0.9.3 leaves in place is the contract underneath: a request +that needs memory an idle job holds is refused once and served on its retry +(itd-2609100457007827's audit; the self-test's watcher), and a model that is +still loading is never evicted at all, whoever holds it. The self-test and the +tool-call probe already yield to a real request through the pool's soft hold +(2026-09-19); the context probe cannot, because it drives the gateway like a +client and a client's hold is never soft (2026-09-11, 2026-09-19). So the +outage's shape — idle work holding a loading model while people wait — is +narrower now but still possible, and the one refusal it costs lands on chat +clients that do not retry. This intent closes that: idle work yields to a real +request, including while its model is loading, and the person asking is parked +for a bounded moment rather than refused. + +## Mechanism + +We expect a real request to be served rather than refused while an idle job +holds the memory because the pool already parks a pre-empting client and tears +down a soft-held ready model for it (the self-test's yield since 2026-09-19), +so the work is to extend two existing seams — the probe takes its hold through +the pool with the same soft-hold tag, and the eviction plan may name a loading +entry whose only holder is soft — rather than to add a scheduler. We are wrong +if a torn-down load leaves the pool's accounting or the model's measurement +inconsistent (a partial figure published, a charge not released, a child not +reaped within the drain bound), or if the park routinely exceeds the drain +bound because a child mid-prefill ignores SIGTERM, in which case the promise +degrades to today's refuse-then-retry. + +## Scope Conditions + +- One serving process per Mac, the pool's memory budget as the only arbiter of what fits; Apple Silicon, the pinned mlx-lm runtime. +- Idle work only: the context probe, the self-test and the tool-call probe are the holders a real request may pre-empt; a hold taken for a client's request is never pre-empted, and no network client can declare or trigger a soft hold. +- A pinned model is never freed for a request, whoever holds it; the client is refused naming the holder. +- The park is bounded by the pool's drain limit (derived from the child stop timeouts, stated in the docs); past it the client is refused naming the holder, not left waiting. +- A yield leaves no partial figure: a yielded probe resumes from its last verified size, a yielded self-test records a yield, not a failure. +- The silent-child readiness residual (a child that answers `/health`, writes no traceback and never answers a completion) is outside this intent — its own capture. + +## Acceptance Criteria + +- **Given** the context probe holding a model, **when** a request needs that + memory, **then** the probe's hold is a soft hold taken through the pool and + the pool pre-empts it exactly as it does the self-test's. + *Held by* a pool test on the probe's hold tag, and an architecture test that + no soft-hold tag is reachable from a gateway request. +- **Given** an idle job's model still loading and a request that needs its + memory, **when** the request arrives, **then** the loading child is torn + down, the request is parked, and it is served without a 503. + *Held by* a pool test driving a slow fake load. +- **Given** a parked request, **when** the teardown exceeds the drain bound, + **then** the request is refused naming the holder and the bound; the bound + is derived from the stop timeouts and stated in `docs/`. + *Held by* a pool test with a child that ignores SIGTERM, and a docs test on + the figure. +- **Given** a pinned model an idle job is using, **when** a request needs its + memory, **then** the pin stays resident and the request is refused naming + the holder; the probe's own unload never unpins. + *Held by* a pool test and a probe test. +- **Given** a probe step or self-test run that is pre-empted, **when** it + yields, **then** no partial figure is published, the probe resumes from its + last verified size, and the self-test records a yield. + *Held by* the existing yield tests, widened to the pre-emption path. +- **Given** a client's own request holding a model, **when** another request + needs the memory, **then** nothing changes from today: a client's hold is + never pre-empted. + *Held by* the existing eviction tests staying green. +- **Given** the panel, **when** an idle job holds a model, **then** the card + names the job and for how long, and Unload frees it as today. + *Held by* a node test over the card. +- **Given** `docs/context-probe.md` and `docs/self-test.md`, **when** read, + **then** they say a real request pre-empts idle work, the park bound, and + that a pin is never freed. + *Held by* the docs-currency hand check on the shipping line. + +## Decisions at the planning interview (2026-09-21) + +Maintainer's answers, each recorded in `.abcd/work/DECISIONS.md`: + +1. **Scope: the full residual.** A real request is served, not refused-then- + retried, while any idle job holds the memory — including while the job's + model is still loading. +2. **Mechanism: a pool-side soft hold for the probe.** The probe holds its + model through the pool with the self-test's soft-hold tag, then drives the + gateway path as before. This supersedes the "keeps internal/runtime + untouched" clause of the 2026-09-11 harness decision; the harness shape + (the gateway path, the gateway's bounds) stands. +3. **A loading model held only by idle work is torn down** for a real request. + This reverses "never evict a model that is still loading" (2026-08-02) for + idle-only holders and for nothing else: a client's load and a pin are + untouched. +4. **The pin wins.** A pinned model is never freed for a request; the client + is refused naming the holder. The probe's own unload path, which ignores + pins today, is a bug captured separately. +5. **The park is bounded by the runtime's drain limit**, derived from the stop + timeouts, never invented; past it the client is refused naming the holder. +6. **The silent-child residual is its own capture**, flagged as reversing the + step-floor decision of 2026-09-21 if its fix needs to. +7. **Impact additive, severity minor**: clients only gain; after 0.9.3 the + residual is one refusal-then-retry. + +Typed links: refines itd-2609091301112705 (answers its open question on +whether the yielding seam becomes real pre-emption); refines +cond-2609100502034205 of itd-2609100457007827 ("the pool has no preemption", +narrowed a second time); supersedes one clause of the 2026-09-11 decision and +one rule of 2026-08-02 as above. + +## Open Questions + +- Whether the pool's drain bound as derived today (about twenty seconds with + eviction grace off) is acceptable as the stated park, or the stop timeouts + are tightened for a soft-held child; the spec's design review proposes, + from measurements on the scratch root. + +## Audit Notes + +_Empty. Populated by intent-auditor when intent moves to shipped/._ + +## Grounds + +- pursued: we expect that once idle work yields to real requests, nobody on the network need know the server measures models in the background; we are wrong if the box's statistics after the run still show refusals attributed to an idle holder, or if pre-emption makes the probe's figures unstable diff --git a/.abcd/development/research/notes/2026-09-06-decomposition-calibration.md b/.abcd/development/research/notes/2026-09-06-decomposition-calibration.md index db0f3410..b893a436 100644 --- a/.abcd/development/research/notes/2026-09-06-decomposition-calibration.md +++ b/.abcd/development/research/notes/2026-09-06-decomposition-calibration.md @@ -713,3 +713,27 @@ the mechanism the design review's most severe finding turned on (the client asks the server about its own tagged request, not SSE comment lines) and sequenced the fair turn first, because its ordering rule is what a place means. Grade: routing survived; one capture landed late. + +## 2026-09-21 — itd-2609211335097114, a stuck model never holds the server hostage + +Filed as a seed from the live-box incident of the same day; two adversarial +reviews (design/feasibility, record discipline) before the interview, both +REVISE on the same ground: most of the seed was already shipped. + +| part | type | home | +|---|---|---| +| A real request pre-empts idle work, including a loading model; the pin wins; a bounded park | capability | itd-2609211335097114 (planned, spc-2609211753023984) | +| A load the child has given up on fails in seconds; a failed model is not re-queued; the probe measures only chat models; the 503 names the holder | already shipped | PR 144, 0.9.3 — struck from the intent, cited in Why This Matters | +| The probe's own unload ignores pins | bug | iss-2609211754251373 | +| A silent child (answers /health, no traceback, no completion) still costs the readiness timeout | bug, reversal-flagged | iss-2609211754253878 | +| Which holders are preemptible; a loading soft-only entry may be torn down; no HTTP client can carry a soft hold | decision (candidate ADR) | the 2026-09-21 decision line (supersedes one clause of 2026-09-11, reverses 2026-08-02 for idle holders); promote to an ADR when the spec's design review confirms the lock order | +| The card's "for how long" | plumbing | criterion 7 of the intent, not its own record | + +Verdict proposed by the reviewers: REVISE — strike the shipped claims, scope +to the residual, name the reversals. Verdict adopted: the full residual, +FILE-AS-REVISED, confirmed by the maintainer at the interview; the reversals +were put to the maintainer as choices (pool-side soft hold; tear down a +loading idle-held model; the pin wins) and each was chosen, not assumed. Two +captures landed at the interview, not late. Grade: routing survived; the +reviewers' "candidate ADR" is carried as a decision line pending the design +review, which is a deferral of the ADR, not a disagreement with the routing. diff --git a/.abcd/development/specs/open/spc-2609211753023984-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md b/.abcd/development/specs/open/spc-2609211753023984-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md new file mode 100644 index 00000000..9964d1a8 --- /dev/null +++ b/.abcd/development/specs/open/spc-2609211753023984-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md @@ -0,0 +1,133 @@ +--- +id: spc-2609211753023984 +slug: a-stuck-model-never-holds-the-server-hostage-alice-asks-her +intent: itd-2609211335097114 +origin: researcher-authored +production_mode: dictated-and-formatted +--- + +# A stuck model never holds the server hostage: a real request pre-empts idle work + +## Summary + +Delivers itd-2609211335097114: a real request is served, not refused and +retried, while an idle job (the context probe, the self-test, the tool-call +probe) holds the memory it needs — including while the job's model is still +loading. Two existing seams are extended: the context probe takes its hold +through the pool with the soft-hold tag the self-test already uses, and the +pool's eviction plan may name a *loading* entry whose only holders are soft. +A pinned model is never freed. The park a pre-empted client pays is bounded +by the pool's drain limit, derived from the child stop timeouts and stated +in the docs; past it the client is refused naming the holder. The maintainer's +seven decisions of 2026-09-21 (on the intent) are the design's fixed points. + +## Scope + +Packages: `internal/runtime` (the pool: eviction plan, soft-only loading +entries, the drain bound as a published figure), `internal/contextprobe` and +`internal/app` (the probe's hold and its yield path; the probe's own unload +respects pins), `internal/gateway` (the refusal after the bound names the +holder — extending `idleHolder`, no new disclosure), `internal/ui` (the card's +"for how long"), `internal/archtest` (no soft-hold tag reachable from a +gateway request), `docs/context-probe.md`, `docs/self-test.md`, and the page +that documents the memory refusal. Changelog under `[Unreleased]`, +`impact: additive`. + +## Approach + +1. **The probe's hold is a pool soft hold** (decision 2). Before a step, the + probe acquires the model through the pool under + `runtime.WithSoftHold(runtime.WithSource(ctx, probeSource), selftest.YieldFrom(ctx))` + — the shape `internal/app/selftest.go:48` and `toolprobe.go:53` use — and + only then drives the gateway path for the step itself, so the measurement + keeps the gateway's bounds (the 2026-09-11 harness shape stands; its + "runtime untouched" clause is superseded by decision 2). The gateway's + own acquisition for the probe's request finds the model resident under the + probe's hold; the probe's hold is the one the pool may pre-empt. The yield + signal reaches the probe as it reaches the self-test (`YieldFrom`); the + probe's `stepYielded` path (`probe.go:329-335`) is unchanged: bounds kept, + no figure published, resume from `b.lo`. +2. **A loading entry with only soft holders is takeable** (decision 3). + `evictionPlanLocked` (`pool.go` ~1624) today skips `!isReady(e)`; it now + admits an entry that is loading AND `softOnlyLocked(e)`, and `preemptLocked` + tears the child down through the ordinary stop path (SIGTERM, SIGKILL, + reap). A client's load and a pin are never in the plan (decision 4; the pin + check precedes the soft-hold clause as today). The abandoned load's charge + is released with the entry; no measurement is written. +3. **The park is the drain bound** (decision 5). A pre-empting caller waits + `DrainWait` (= `maxDrainWait` = `stopBound + drainMargin`, `pool.go:797-801`) + for the memory to come back, then is refused. The figure is exported once + (`runtime.DrainBound()` or the pool's option, one canonical place) and the + docs state it from that source; the refusal after the bound extends the + 0.9.3 `idleHolder` text with the bound — to entitled clients only, as the + 2026-09-21 disclosure rule already limits it. +4. **The pin wins** (decision 4). The probe's `unloadWaiting` → `Pool.Unload` + path gains the pin check the eviction plan has (the bug captured beside + this spec): a probe never unloads a pinned model on a yield; the client is + refused naming the holder. +5. **The card says for how long** (criterion 7): `residencyLabel` in + `internal/ui/static/app.js` reads `selftest.Status.Since` (already on the + snapshot) and the probe's equivalent, and renders the duration beside the + job. +6. **Docs** (criterion 8): `docs/context-probe.md` and `docs/self-test.md` + say a real request pre-empts idle work, the park bound and its source, and + that a pinned model is never freed; the memory-refusal page carries the new + sentence. + +## How each criterion is held + +1. Probe's hold is a pool soft hold — a pool test asserting the probe's + acquisition is `softOnlyLocked`; an archtest walking the gateway's request + path for any `WithSoftHold` (none reachable). +2. Loading model torn down, request served — a pool test with a fake launcher + whose load blocks: a soft-held loading entry, a client acquisition, assert + the child stopped, the client served, no 503. +3. The park bound — a pool test with a fake child that ignores SIGTERM: the + client is refused after `DrainWait` with the holder and the bound in the + error; a docs test holding the documented figure to the exported constant. +4. The pin wins — a pool test (pinned soft-held entry: not in the plan, client + refused naming the holder) and a probe test (`unloadWaiting` on a pinned + model does not unload). +5. A yield leaves no partial figure — the existing `stepYielded` and self-test + yield tests, run through the new pre-emption path (loading and ready). +6. A client's hold is never pre-empted — the existing eviction tests stay + green; one added: a client-held loading entry is not in the plan. +7. The card — a node test over `residencyLabel` with `Since` set. +8. The docs — the docs-currency hand check on the shipping line. + +## Security + +`internal/runtime` and `internal/gateway` are trust boundaries. The soft-hold +tag stays a context value no HTTP request can set (archtest, criterion 1). +Tearing down a loading child is the existing stop path; nothing new is +executed. The refusal names the holder only to entitled clients, as 0.9.3 +does; the bound is a constant, not a secret. Adversarial security review of +the diff before it is presented. + +## Verification + +- `make test` green, `gofmt -l .` empty, `go vet ./...` clean; every test + above watched red before the change and green after. +- Hand checks, recorded on the shipping decision line, on the scratch root + with the probe on and the idle threshold low: send a chat request while the + probe is loading its model and see it answered (the log shows the yield, + the park and the load); pin the model under probe and see the refusal name + the holder with the pin resident; read the park bound off the docs and off + the refusal and see them agree; the card's duration. +- The open question (whether about twenty seconds is acceptable as the park, + or the stop timeouts are tightened for a soft-held child) is answered by + the design review from measurements on the scratch root, and the answer + recorded as a departure or a decision line. + +## Out of scope + +The silent-child readiness residual (its own capture); any change to what +the probe charges; a new setting or a config.json change; pre-empting a +client's hold; the served-window proposal (itd-2609091712142715, held). + +## Falsifiers + +The mechanism claim on the intent: a torn-down load that leaves the +accounting or a measurement inconsistent, or a park that routinely exceeds +the drain bound. Either degrades the promise to refuse-then-retry and is +recorded as a departure, not hidden. diff --git a/.abcd/work/DECISIONS.md b/.abcd/work/DECISIONS.md index fa29c0df..f7228058 100644 --- a/.abcd/work/DECISIONS.md +++ b/.abcd/work/DECISIONS.md @@ -386,3 +386,5 @@ - 2026-09-21 — **The stuck context probe is three bug fixes, and pre-emption stays a draft** (branch `fix/stuck-probe`; iss-2609211334563318, iss-2609211334570516 and iss-2609211334576018 resolved; itd-2609211335097114 untouched). The live server (v0.9.1) refused every chat request 503 for six hours because the probe queued `mlx-community/GLM-OCR-bf16` — an image-to-text model, `chat: false` — and loaded it thirty-two times: the child raised `ValueError: Model type glm_ocr not supported` in its generate thread on the first request while its httpd answered `/health`, so the pool waited its ten-minute readiness timeout each time, and while loading the model was charged the whole budget (its default served window is worked out to fill what the budget has, so a 2.2 GB model with no served-window setting is charged ≈ the budget from the moment its entry exists — a property of the charge, not of loading, and not changed here). (1) The probe considers only models the server offers to chat: `Candidates()` and `MeasureNow` read `registry.Model.CanChat` with the rule in force. (2) A load the child has given up on fails in seconds: the pool watches the per-model child log while it waits for readiness (`LoadLogger`, which the real launcher's process satisfies) and ends the wait on a traceback whose terminal line is a ValueError, ModuleNotFoundError or ImportError, the line bounded and stripped of anything path-shaped; BrokenPipeError and the like are not in the set because the child survives them. A load that never became ready is recorded on the model (`registry.LoadFailure`: reason and provenance — runtime, budget, concurrency, served window); while it stands the probe and the self-test skip the model, a queued probe of it is dropped, and a request for it is refused at once as a NotReadyError carrying the reason; it is lifted by a moved provenance (through `RefreshStaleness`), a re-download, Load or Measure now. The pool tells the observer whether a failure was the model's own or interrupted (the entry taken out of the pool meanwhile), and only the former is recorded. The probe's ten-minute step floor is deliberately left: it is the gateway's own prefill base, which the probe's timer must not undercut or a slow step is filed as the deadline's, and a step's request includes the cold load the pool allows ten minutes for. (3) The refusal names the holder to an entitled client only: the gateway is handed the idle loop's status (which gains `since`), and a no-room refusal to a loopback or key-admitted client — the same clients the models list tells what is resident — names the model the job holds, the job, for how long, and that Unload releases it; the pool's own refusal still names no model and unentitled clients still get the generic sentence. The card's pill says "loading", "loading for the context probe", "held by the self-test". What would show these wrong: a chat model the probe now skips; a genuine load — a slow cold load — that a traceback line in the set fails early; a load failure that survives a runtime change; a keyless network client that reads a model id in a 503. Pre-emption — a real request taking the memory idle work holds — is itd-2609211335097114, a draft with no acceptance criteria, and is not implemented or approximated here. - 2026-09-21 — **Corrections to the stuck-probe line above, from its two adversarial reviews** (an append-only ledger corrects by superseding; both lines stand and this one governs where they differ). (1) A load failure that is the pool's own bound — the readiness timeout, an exit by signal — is recorded as `transient`: it stands for this process (idle work skips the model, a client is told why at once) and is dropped at the next start, because a slow load on a busy Mac says nothing about the next one; the child's own traceback and a non-signal exit status stand until the provenance moves or a person retries, as the line above says. (2) The refusal's promise that Unload releases a model an idle job holds was false while the job's own request was in flight (the pool refused it as busy — the maintainer's two 409s on the live box): the idle loop now exposes `Runner.Interrupt(model)`, the run yields as it does for a client, and the panel's Unload asks for that first and waits, bounded, for the model to go (`TestUnloadFromThePanelTakesTheModelBackFromTheProbe`, watched red on the 409). (3) The child-log reader takes a terminal line only straight after a traceback's frames and only from whole lines; its open refuses a link and a FIFO; the sanitiser drops control characters and blanks a path with spaces as one path. (4) A failure reported after a hand retry has started a fresh load is not written over it; a Measure now arriving between Due's candidate snapshot and its pruning is not pruned. Accepted without change: a local process that can reach the child's loopback port can write a line the reader takes (the same trust class that can already plant registry.json; no new privilege); the provenance is read when the failure is recorded rather than when the load began. - 2026-09-21 — **The 0.9.3 cut, codename Prellerhaus**, by hand on the 0.9.2 precedent (PR 140), on the maintainer's decision of 2026-09-21. The cut ships the stuck-probe fixes (iss-2609211334563318, iss-2609211334570516, iss-2609211334576018; PR 144, `impact: fix`) and the no-transcript models intent already on `main` without a release (itd-2609091715089488, PR 141, `impact: additive`), so the version is a patch; `build/CODENAME` is unchanged. Three open majors are deferred out loud with `deferred_after: "v0.9.3"`: iss-2609211218478273, because the no-transcript server half is live and every model truthfully reads "keeps no transcript" until the parent (itd-2609091707499248) lands, which is the next lane; iss-2609200815308397 and iss-2609190242198542 re-deferred with their existing reasons, because the maintainer tests by hand against this release and neither is of this cut's class. `abcd launch ship` refused the cut twice — the surface guard has no baseline in v0.9.2's tree (seeded after it; the clean-cutover manual roll is what puts one into a tag), and the unfixed-finding gate reads a deferral as `deferred_after: ` where this repository's records write the cut version — so the roll is manual, as the precedent's was; both are tooling findings for abcd, not this ledger. Retention after verification deletes the v0.9.2 release and keeps its tag. +- 2026-09-21 — **itd-2609211335097114 a stuck model never holds the server hostage: planned as the full residual — a real request pre-empts idle work, including while the job's model is still loading; the probe's hold becomes a pool-side soft hold; a loading model held only by idle work is torn down; the pin wins; the park is the runtime's drain bound; impact additive, severity minor** (maintainer's decisions at the 2026-09-21 planning interview, after two adversarial reviews — design/feasibility and record discipline — both REVISE: most of the seed was already shipped in 0.9.3 by PR 144 and by the 2026-09-19 soft hold, and the residual touches four recorded decisions). (1) Scope: the full residual, not the probe's hold alone and not withdrawal. (2) Mechanism: the context probe acquires its model through the pool under the self-test's soft-hold tag and then drives the gateway path as before — this SUPERSEDES the "keeps internal/runtime untouched" clause of the 2026-09-11 harness decision; the harness shape (the gateway path and its bounds) stands. (3) A loading entry whose only holders are soft may be named by the eviction plan and torn down for a real request — this REVERSES "never evict a model that is still loading" (2026-08-02) for idle-only holders and for nothing else: a client's load and a pin are untouched. (4) A pinned model is never freed for a request, whoever holds it; the client is refused naming the holder; the probe's own unload path ignoring pins is a bug captured separately. (5) The park is bounded by the pool's drain limit, derived from the child stop timeouts and stated in the docs from one exported figure, never invented; past it the client is refused naming the holder. (6) The silent-child readiness residual (answers `/health`, no traceback, no completion) is its own capture, flagged against the step-floor decision of the same day. (7) Impact additive (clients only gain; the pin and client-hold contracts are held by tests), severity minor (after 0.9.3 the residual is one refusal-then-retry). Typed links: refines itd-2609091301112705 (answers its open question on real pre-emption), refines cond-2609100502034205 of itd-2609100457007827 (narrowed a second time). Spec spc-2609211753023984; grounds recorded on the intent; READY. +- 2026-09-21 — **itd-2609091712142715 (the served window proposes itself) stays held; the hold is lifted for evidence-gathering only** (maintainer, at the same interview). The 2026-09-20 hold stands and is lifted by evidence, not by decision; what changes is that the evidence is now sought on purpose: the autonomous run of 2026-09-23 carries a task to turn the self-test on for one cycle on the live box (v0.9.3, statistics on since the restart of 2026-09-21) and to read back the served-window statistics — the prompt-size distribution against the served window per model, the served-window refusal rate, the observed in-flight counts and the sampling override rate — onto the record, which is exactly the reading the hold names. The interview for that draft happens once the reading exists. diff --git a/.abcd/work/issues/open/iss-2609211754251373-the-context-probe-s-own-unload-path-ignores-pins-on-a-yield.md b/.abcd/work/issues/open/iss-2609211754251373-the-context-probe-s-own-unload-path-ignores-pins-on-a-yield.md new file mode 100644 index 00000000..fb31ca71 --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211754251373-the-context-probe-s-own-unload-path-ignores-pins-on-a-yield.md @@ -0,0 +1,14 @@ +--- +schema_version: 1 +id: "iss-2609211754251373" +slug: "the-context-probe-s-own-unload-path-ignores-pins-on-a-yield" +severity: "minor" +category: "bug" +source: "plan-review" +found_during: "planning interview itd-2609211335097114 (design review finding 7)" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/contextprobe/probe.go" +--- + +The context probe's own unload path ignores pins: on a yield or a held step the probe calls unloadWaiting -> Pool.Unload (internal/contextprobe/probe.go), which unloads whatever it is given, while the eviction plan skips a pinned model before it looks at soft holds ('a pin still beats everything', 2026-09-19) and the probe intent says it never evicts a pinned model. So a pinned model under probe is freed by one path and protected by the other. The pin must win on both paths (decision 4 of itd-2609211335097114's interview); the fix is the probe's, not the pool's. diff --git a/.abcd/work/issues/open/iss-2609211754253878-the-readiness-wait-cannot-tell-an-idle-but-silent-child-from.md b/.abcd/work/issues/open/iss-2609211754253878-the-readiness-wait-cannot-tell-an-idle-but-silent-child-from.md new file mode 100644 index 00000000..f69ff460 --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211754253878-the-readiness-wait-cannot-tell-an-idle-but-silent-child-from.md @@ -0,0 +1,14 @@ +--- +schema_version: 1 +id: "iss-2609211754253878" +slug: "the-readiness-wait-cannot-tell-an-idle-but-silent-child-from" +severity: "minor" +category: "bug" +source: "plan-review" +found_during: "planning interview itd-2609211335097114" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/runtime/launcher.go" +--- + +The readiness wait cannot tell an idle-but-silent child from a slow one: a model server that answers /health but never answers a completion and writes no traceback (no fatal line for the LoadLogger of PR 144 to act on) still costs the full ten-minute readiness timeout, once per process (marked transient). Seen only in principle; the 2026-09-21 incident's child DID write a traceback. The fix would detect the silent case (e.g. a bounded first completion during readiness) and, if it lowers the probe's step floor to do so, REVERSES the 2026-09-21 decision that keeps the floor as the gateway's prefill base — flag for the maintainer, never assert. Routed out of itd-2609211335097114 at its planning interview (decision 6). diff --git a/.abcd/work/issues/open/iss-2609211754256577-evidence-gathering-task-for-the-autonomous-run-of-2026-09-23.md b/.abcd/work/issues/open/iss-2609211754256577-evidence-gathering-task-for-the-autonomous-run-of-2026-09-23.md new file mode 100644 index 00000000..b1633aac --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211754256577-evidence-gathering-task-for-the-autonomous-run-of-2026-09-23.md @@ -0,0 +1,13 @@ +--- +schema_version: 1 +id: "iss-2609211754256577" +slug: "evidence-gathering-task-for-the-autonomous-run-of-2026-09-23" +severity: "minor" +category: "process" +source: "plan-review" +found_during: "planning interview 2026-09-21" +origin: researcher-authored +production_mode: hand-written +--- + +Evidence-gathering task for the autonomous run of 2026-09-23, per the 2026-09-21 decision on itd-2609091712142715: with the live box on v0.9.3 and statistics on, run ONE self-test cycle and read back the served-window statistics onto the record (prompt-size distribution against the served window per model, served-window refusal rate, observed in-flight counts, sampling override rate per parameter) as a dated note under .abcd/development/research/ — the reading the hold on itd-2609091712142715 names. Not a code change; the run's script carries it as a manual/handcheck-class step because it touches the live box, which no session contacts on its own.