From 335b1a78630b21af2dc191d3858cce8ab7e7cb2c Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Mon, 21 Sep 2026 16:01:38 +0100 Subject: [PATCH] chore: capture the stuck context probe and draft the pre-emption intent Live server 2026-09-21, six hours of 503s: the context probe queued an image-to-text model mlx_lm cannot load (Model type glm_ocr not supported), waited the ten-minute readiness timeout, marked it load_failed, re-queued it and retried 32 times, each attempt holding the loading model charged at the whole budget so no client request fit. Three issues and one draft intent (a real request pre-empts idle work). Assisted-by: Claude Opus 5 (claude-opus-5) --- ...holds-the-server-hostage-alice-asks-her.md | 43 +++++++++++++++++++ ...sures-models-that-are-not-chat-models-i.md | 14 ++++++ ...be-calibrates-a-model-the-model-is-load.md | 34 +++++++++++++++ ...s-while-the-budget-is-held-by-idle-work.md | 14 ++++++ 4 files changed, 105 insertions(+) create mode 100644 .abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md create mode 100644 .abcd/work/issues/open/iss-2609211334563318-the-context-probe-measures-models-that-are-not-chat-models-i.md create mode 100644 .abcd/work/issues/open/iss-2609211334570516-while-the-context-probe-calibrates-a-model-the-model-is-load.md create mode 100644 .abcd/work/issues/open/iss-2609211334576018-the-503-a-client-gets-while-the-budget-is-held-by-idle-work.md diff --git a/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md b/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md new file mode 100644 index 00000000..2756a638 --- /dev/null +++ b/.abcd/development/intents/drafts/itd-2609211335097114-a-stuck-model-never-holds-the-server-hostage-alice-asks-her.md @@ -0,0 +1,43 @@ +--- +id: itd-2609211335097114 +slug: a-stuck-model-never-holds-the-server-hostage-alice-asks-her +spec_id: null +kind: null +suggested_kind: null +reclassification_history: [] +builds_on: [] +severity: minor +impact: additive +origin: researcher-authored +production_mode: hand-written +--- + +# A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed. + +## Press Release + +> _Seeded from a quoted-text intent capture. Expand into the full press-release narrative before planning._ + +## Why This Matters + +A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed. + +## Mechanism + +> _Prompted (the claim-recording gradient): why the authors expect this to work, as a falsifiable "we expect X because Y" — not the outcome restated. Replace this line with the claim, or with the exact token `None stated.` alone on its line to record the claim as considered and declined._ + +## Scope Conditions + +> _Required (the claim-recording gradient): the population, platform, scale, or assumptions this claim holds under, one per top-level bullet — `abcd intent plan` stamps each with a persistent identity. Replace this line with those bullets, or with the exact token `None stated.` alone on its line._ + +## Acceptance Criteria + +> _Required (the itd-1 discipline): add at least one Given-When-Then bullet describing the verifiable bar for "shipped" before this draft can be planned._ + +## Open Questions + +_None recorded yet._ + +## Audit Notes + +_Empty. Populated by intent-auditor when intent moves to shipped/._ diff --git a/.abcd/work/issues/open/iss-2609211334563318-the-context-probe-measures-models-that-are-not-chat-models-i.md b/.abcd/work/issues/open/iss-2609211334563318-the-context-probe-measures-models-that-are-not-chat-models-i.md new file mode 100644 index 00000000..789bc58a --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211334563318-the-context-probe-measures-models-that-are-not-chat-models-i.md @@ -0,0 +1,14 @@ +--- +schema_version: 1 +id: "iss-2609211334563318" +slug: "the-context-probe-measures-models-that-are-not-chat-models-i" +severity: "major" +category: "bug" +source: "user-observation" +found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/app/contextprobe.go" +--- + +The context probe measures models that are not chat models: it takes every Ready() model as a candidate (internal/app/contextprobe.go Candidates), so on the live server it picked mlx-community/GLM-OCR-bf16 (pipeline image-to-text, chat: false in the models list), spawned an mlx_lm server for it and POSTed /v1/chat/completions, which never answered (the child idle at 0% CPU answering /health in under a millisecond, in_flight 1 for over ten minutes at 'calibrating at 1024 tokens'). A served window is only meaningful for a model the server offers to chat; the probe should skip chat: false models, and a step whose child answers /health but not a completion should fail fast rather than wait for the step timeout. diff --git a/.abcd/work/issues/open/iss-2609211334570516-while-the-context-probe-calibrates-a-model-the-model-is-load.md b/.abcd/work/issues/open/iss-2609211334570516-while-the-context-probe-calibrates-a-model-the-model-is-load.md new file mode 100644 index 00000000..59527566 --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211334570516-while-the-context-probe-calibrates-a-model-the-model-is-load.md @@ -0,0 +1,34 @@ +--- +schema_version: 1 +id: "iss-2609211334570516" +slug: "while-the-context-probe-calibrates-a-model-the-model-is-load" +severity: "major" +category: "bug" +source: "user-observation" +found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/contextprobe/probe.go" +--- + +While the context probe calibrates a model, the model is loading/in_flight and charged the whole memory budget (charge_bytes = budget: 82.46 GB for a 2.2 GB model), so it is not evictable and every real request for another model is refused 503 for the length of the step; the step timeout floor is eleven minutes (defaultStepTimeout in internal/contextprobe/probe.go: at least ten minutes plus a minute's margin), and after a failed step the probe can pick the same model again. Idle work starved real requests on the live server for the whole period the maintainer tried several chat clients. An idle job should yield the budget to a real request (abort the step, release the model) rather than the request yielding to the job. + +- 2026-09-21 14:35 (observed by the operator, not inferred): the probe re-picked the same model after the step ended — `loaded_at` moved from 14:26:50 to 14:34:35 with the step still "calibrating at 1024 tokens" — so it loops on the unresponsive model and the budget is never released. + +## Correction from the live box's logs (2026-09-21 15:55, `dessau.log` and the child's log) + +The mechanism is the LOAD path, not the probe's step timeout. The child +(`mlx_lm server`) raised `ValueError: Model type glm_ocr not supported` in its +generate thread on the first request while its httpd kept answering `/health`; +Dessau waited its load-readiness timeout — "did not become ready within 10m0s" +— then logged `model failed to load`, unloaded it (`reason=load_failed`), and +the probe re-queued it ("the measurement is held … it stays queued") and tried +again eleven minutes later: **32 attempts from 09:36 to 15:38**, six hours, +each holding the loading model charged at ≈ the whole budget so no client +request fit (the clients' 503s took 0–1 ms; 32 of the 35 logged 503s are the +probe's own requests after 10 min). Only the maintainer's Unload at 15:49 made +the probe record the failure (`bound=model`) and stop. So the fixes are: (1) a +load fails fast when the child's own log carries a fatal load error (Dessau +already captures that log per model) instead of waiting the readiness timeout; +(2) a `load_failed` model is not re-queued by an idle job until something about +it changes; (3) the probe's step floor is secondary and stays as filed. diff --git a/.abcd/work/issues/open/iss-2609211334576018-the-503-a-client-gets-while-the-budget-is-held-by-idle-work.md b/.abcd/work/issues/open/iss-2609211334576018-the-503-a-client-gets-while-the-budget-is-held-by-idle-work.md new file mode 100644 index 00000000..5ee3feb7 --- /dev/null +++ b/.abcd/work/issues/open/iss-2609211334576018-the-503-a-client-gets-while-the-budget-is-held-by-idle-work.md @@ -0,0 +1,14 @@ +--- +schema_version: 1 +id: "iss-2609211334576018" +slug: "the-503-a-client-gets-while-the-budget-is-held-by-idle-work" +severity: "minor" +category: "ux" +source: "user-observation" +found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/gateway" +--- + +The 503 a client gets while the budget is held by idle work reads 'not enough memory to load another model, and no model in memory can be freed (limit 76.8 GB)': true, but it names neither what holds the memory (a loading model under the context probe) nor that the holder is idle work the server started itself, so the person reads it as their model being too big. The refusal should name the holder and the job, and the panel should show the same (the card says only 'loading').