Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
id: itd-2609211335097114
slug: a-stuck-model-never-holds-the-server-hostage-alice-asks-her
spec_id: null
kind: null
suggested_kind: null
reclassification_history: []
builds_on: []
severity: minor
impact: additive
origin: researcher-authored
production_mode: hand-written
---

# A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed.

## Press Release

> _Seeded from a quoted-text intent capture. Expand into the full press-release narrative before planning._

## Why This Matters

A stuck model never holds the server hostage. Alice asks her Dessau Server a question and gets an answer, even while the server is busy with work it started on its own: an idle job such as the context probe yields the moment a real request needs the memory it holds — the step is abandoned, the model it was measuring is released, and the request loads its model as if nothing had been in the way. When something does hold the memory, Alice can see it: the control panel names the model, the job holding it and for how long, and offers one click to free it; a client refused for lack of memory is told what holds it and that it is the server's own work, not the size of the model asked for. A job that finds its model unresponsive gives up in seconds, not after a timeout sized for a hundred thousand tokens, and does not pick the same model again until something about it has changed.

## Mechanism

> _Prompted (the claim-recording gradient): why the authors expect this to work, as a falsifiable "we expect X because Y" — not the outcome restated. Replace this line with the claim, or with the exact token `None stated.` alone on its line to record the claim as considered and declined._

## Scope Conditions

> _Required (the claim-recording gradient): the population, platform, scale, or assumptions this claim holds under, one per top-level bullet — `abcd intent plan` stamps each with a persistent identity. Replace this line with those bullets, or with the exact token `None stated.` alone on its line._

## Acceptance Criteria

> _Required (the itd-1 discipline): add at least one Given-When-Then bullet describing the verifiable bar for "shipped" before this draft can be planned._

## Open Questions

_None recorded yet._

## Audit Notes

_Empty. Populated by intent-auditor when intent moves to shipped/._
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
schema_version: 1
id: "iss-2609211334563318"
slug: "the-context-probe-measures-models-that-are-not-chat-models-i"
severity: "major"
category: "bug"
source: "user-observation"
found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request"
origin: researcher-authored
production_mode: hand-written
found_at: "internal/app/contextprobe.go"
---

The context probe measures models that are not chat models: it takes every Ready() model as a candidate (internal/app/contextprobe.go Candidates), so on the live server it picked mlx-community/GLM-OCR-bf16 (pipeline image-to-text, chat: false in the models list), spawned an mlx_lm server for it and POSTed /v1/chat/completions, which never answered (the child idle at 0% CPU answering /health in under a millisecond, in_flight 1 for over ten minutes at 'calibrating at 1024 tokens'). A served window is only meaningful for a model the server offers to chat; the probe should skip chat: false models, and a step whose child answers /health but not a completion should fail fast rather than wait for the step timeout.
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
---
schema_version: 1
id: "iss-2609211334570516"
slug: "while-the-context-probe-calibrates-a-model-the-model-is-load"
severity: "major"
category: "bug"
source: "user-observation"
found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request"
origin: researcher-authored
production_mode: hand-written
found_at: "internal/contextprobe/probe.go"
---

While the context probe calibrates a model, the model is loading/in_flight and charged the whole memory budget (charge_bytes = budget: 82.46 GB for a 2.2 GB model), so it is not evictable and every real request for another model is refused 503 for the length of the step; the step timeout floor is eleven minutes (defaultStepTimeout in internal/contextprobe/probe.go: at least ten minutes plus a minute's margin), and after a failed step the probe can pick the same model again. Idle work starved real requests on the live server for the whole period the maintainer tried several chat clients. An idle job should yield the budget to a real request (abort the step, release the model) rather than the request yielding to the job.

- 2026-09-21 14:35 (observed by the operator, not inferred): the probe re-picked the same model after the step ended — `loaded_at` moved from 14:26:50 to 14:34:35 with the step still "calibrating at 1024 tokens" — so it loops on the unresponsive model and the budget is never released.

## Correction from the live box's logs (2026-09-21 15:55, `dessau.log` and the child's log)

The mechanism is the LOAD path, not the probe's step timeout. The child
(`mlx_lm server`) raised `ValueError: Model type glm_ocr not supported` in its
generate thread on the first request while its httpd kept answering `/health`;
Dessau waited its load-readiness timeout — "did not become ready within 10m0s"
— then logged `model failed to load`, unloaded it (`reason=load_failed`), and
the probe re-queued it ("the measurement is held … it stays queued") and tried
again eleven minutes later: **32 attempts from 09:36 to 15:38**, six hours,
each holding the loading model charged at ≈ the whole budget so no client
request fit (the clients' 503s took 0–1 ms; 32 of the 35 logged 503s are the
probe's own requests after 10 min). Only the maintainer's Unload at 15:49 made
the probe record the failure (`bound=model`) and stop. So the fixes are: (1) a
load fails fast when the child's own log carries a fatal load error (Dessau
already captures that log per model) instead of waiting the readiness timeout;
(2) a `load_failed` model is not re-queued by an idle job until something about
it changes; (3) the probe's step floor is secondary and stays as filed.
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
schema_version: 1
id: "iss-2609211334576018"
slug: "the-503-a-client-gets-while-the-budget-is-held-by-idle-work"
severity: "minor"
category: "ux"
source: "user-observation"
found_during: "live server v0.9.1 on 2026-09-21, 503 for every chat request"
origin: researcher-authored
production_mode: hand-written
found_at: "internal/gateway"
---

The 503 a client gets while the budget is held by idle work reads 'not enough memory to load another model, and no model in memory can be freed (limit 76.8 GB)': true, but it names neither what holds the memory (a loading model under the context probe) nor that the holder is idle work the server started itself, so the person reads it as their model being too big. The refusal should name the holder and the job, and the panel should show the same (the card says only 'loading').
Loading