Skip to content

feat(seats): mimo-9b-agent becomes amd-gcn's bound agent seat - #474

Merged
dmmdea merged 3 commits into
mainfrom
feat/mimo-agent-seat-amd-gcn
Sep 24, 2026
Merged

dmmdea merged 3 commits into
mainfrom
feat/mimo-agent-seat-amd-gcn

Conversation

@dmmdea

@dmmdea dmmdea commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Summary

  • amd-gcn (Ryzen 5 5625U / Vega 7 iGPU, Vulkan) now sets include_mimo_9b: true and binds config_seed.agent_model to mimo-9b-agent, keeping include_qwen35_4b: true as the rollback seat. MEASURED 2026-09-24 on the amd-gcn reference box (llama.cpp b11153 Vulkan): mimo-9b-agent ties qwen3.5-4b-agent on quality (shape B 18/18 vs 17/18, shape C 4/5 vs 5/5) at roughly half the wall (333s vs 614s).
  • servingtmpl.Params.validate() no longer refuses IncludeMimo9B && IncludeQ354B. A new __Q354B_AGENT_ALIAS__ token (mirroring the existing __Q359B_AGENT_ALIAS__) drops qwen3.5-4b-agent's own claim on agent-seat when mimo-9b-agent also renders, so the 4B entry stays as an un-aliased rollback instead of being refused. The pre-existing include_qwen35_4b/include_qwen35_9b mutual exclusion is untouched.
  • llama-swap.linux-vulkan.yaml's mimo-9b-agent block now serves the tier's own window/KV/flash-attn/cache-ram (__CTX__/__KV_K__/__KV_V__/__FLASH_ATTN__/__CACHE_RAM__) instead of the CUDA 8GB tiers' literal 65536/q8_0 pin — a Vulkan UMA box has no fixed 8GB-class fit to pin to. win-cuda.yaml and linux-cuda.yaml's mimo blocks are unchanged (still the literal CUDA pin).
  • Not originally scoped, required by two existing gates: llama-swap.win-vulkan.yaml gained a mimo-9b-agent entry (same tier-token shape) it never had. TestEveryBoundAliasIsServed and TestVulkanAndCPUTiersRenderOnLinux both assert every vulkan-backend tier renders on BOTH operating systems ("a tier is a hardware class; the OS it boots must not decide whether it exists"). Setting include_mimo_9b on amd-gcn without this addition made the Windows render of the tier refuse outright (win-vulkan.yaml had no mimo-9b-agent entry), which is exactly the OS-decides-capability failure those two gates exist to catch.
  • REQUIRES llama.cpp >= b11102 (upstream llama.cpp #29319) — documented in the template comments, the tier notes, and docs/tiers/amd-gcn.md.

Tiers that render linux-vulkan.yaml / win-vulkan.yaml

Three tiers share backend: "vulkan": amd-gcn, amd-rdna3, amd-rdna3-dgpu.

  • amd-gcn: the only one affected — gains the mimo-9b-agent seat as described above.
  • amd-rdna3 / amd-rdna3-dgpu: functionally unchanged. Neither sets include_mimo_9b, so the new mimo block strips out of their renders on both OSes exactly as before (no new model, no alias, no matrix membership). Byte-diffed both tiers on both OSes before/after: the only differences are the provenance stamp (harness version bump) and inert prose comment lines preceding the stripped block — the same "comment survives a dropped model" characteristic every other gated seat in these templates already has (verified against the pre-existing mimo comment in win-cuda.yaml, which leaks identically for e.g. blackwell-16).

Rendered output (verification)

go run . install render --profile amd-gcn --os linux --home /opt/offload --llama-bin /opt/llama-vulkan --models /opt/models --threads 6 --listen 127.0.0.1:11436 --root .

mimo-9b-agent:

mimo-9b-agent:
  aliases: [mimo-9b, agent-seat]
  env: ["${ld}", "${vk}"]
  cmd: >-
    /opt/llama-vulkan/llama-server --model /opt/models/MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf
    --n-gpu-layers 999 --parallel 1 --ctx-size 32768 --flash-attn on
    --cache-type-k f16 --cache-type-v f16 --cache-ram 8192  --threads 6
    --reasoning off --jinja --port ${PORT} --host 127.0.0.1
  checkEndpoint: /health
  ttl: 300

qwen3.5-4b-agent (rollback, alias dropped):

qwen3.5-4b-agent:
  aliases: [qwen35-4b]
  env: ["${ld}", "${vk}"]
  cmd: >-
    /opt/llama-vulkan/llama-server --model /opt/models/Qwen3.5-4B-UD-Q4_K_XL.gguf
    --n-gpu-layers 999 --parallel 1 --ctx-size 32768 --flash-attn on
    --cache-type-k f16 --cache-type-v f16 --cache-ram 8192  --threads 6
    --jinja --port ${PORT} --host 127.0.0.1
  checkEndpoint: /health
  ttl: 300

(The double space before --threads is the empty slot-save token, [slot-save token output] in the task's expected shape — unchanged, unrelated to this PR.)

Tests

  • go test ./...: all green (one unrelated pre-existing flake, TestDrainWaitsForARegisteredRunAcrossTheStepGap in gpu_drain_runs_test.go, a live GPU-lease timing test under full-suite load — passes 3/3 in isolation, file untouched by this PR).
  • go vet ./...: clean.
  • setup/render.tests.ps1: ALL PASS (new amd-gcn assertions for the mimo seat, the alias handoff, and the tier-driven ctx/KV).
  • setup/tests/install-config-seed.test.ps1: ALL PASS (replaced the table-wide include_qwen35_4b/include_mimo_9b mutual-exclusion assertion — the exact rule this PR changes — with amd-gcn-specific assertions mirroring the existing blackwell-8/ampere-8 ones).

New/changed tests, each proven able to fail by reverting its production line:

  • TestMimoAndQwen354BRenderTogetherQwenLosesTheAlias (replaces TestMimoAndQwen354BAreRefused) — red when validate()'s old refusal is restored.
  • TestMimo9BSeatKeepsItsMeasuredInvariants (linux-vulkan/win-vulkan branch) — red when the mimo block is reverted to the CUDA literal pin.
  • TestAmdGcnSeedsTheSeatItWasMeasuredOn (Go) and the new amd-gcn assertions in both .ps1 suites — red when agent_model/include_mimo_9b are reverted.
  • TestAgentWindowMatchesWhatTheAgentSeatServes's backend now resolves a tier's agent-seat ctx from the template it actually renders (linux-vulkan for vulkan, win-cuda otherwise) instead of always reading win-cuda.yaml — a latent bug the vulkan+mimo combination exposed (win-cuda's mimo entry is a literal 65536 while amd-gcn's own entry is __CTX__/32768).

Not done / findings

  • install.sh (the Linux install path) does no model downloads at all by design (documented as an operator prerequisite in its header) — there is no separate Linux download gate to wire up. install.ps1's include_mimo_9b gate is purely profile-id-keyed, so amd-gcn triggers it automatically with no code change.
  • agent_seat_tok_s (10) stays as measured for the previous qwen3.5-4b-agent seat — not touched, since the task's config_seed instructions named only agent_model. Flagging in case the operator wants a mimo-specific rate re-measured for contract-wall sizing.
  • Two Vulkan-only templates (win-vulkan.yaml, linux-vulkan.yaml) also carry a __Q354B_ALT__/win-vulkan.yaml gap in q354b_deadtoken_test.go's pre-existing tokenTemplates contract map (it only lists linux-cuda/win-cuda for that token, missing both vulkan templates) — pre-existing, not touched by this PR, noted for a future pass.

🤖 Generated with Claude Code

dmmdea and others added 3 commits September 24, 2026 06:35
…o pairing renders instead of refusing

amd-gcn (Ryzen 5 5625U / Vega 7 iGPU, Vulkan) now sets include_mimo_9b: true and binds
config_seed.agent_model to mimo-9b-agent, keeping qwen3.5-4b-agent as the un-aliased
ROLLBACK seat instead of stripping it — the same shape PR #472 already gave the
mimo/qwen3.5-9b-agent pair, extended to the 4B one weight class down via a new
__Q354B_AGENT_ALIAS__ token. On the Vulkan UMA templates the mimo-9b-agent entry now
follows the tier's own window/KV/flash-attn instead of the CUDA 8GB tiers' literal
65536/q8_0 pin, since a Vulkan UMA box has no fixed 8GB-class fit to pin to.

win-vulkan.yaml gained a mimo-9b-agent entry it never had (OS parity: two existing
gates require every vulkan-backend tier to render on both operating systems, and
amd-gcn's Windows render would otherwise refuse outright). amd-rdna3/amd-rdna3-dgpu,
the other two vulkan tiers, are functionally unaffected (the block strips out of
their renders as before; only inert prose comments differ, matching an existing
pattern already shipped for every other gated seat in these templates).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…w binds

agent_seat_tok_s sizes agent walls until the seat records its own rate. 10 was
qwen3.5-4b-agent's; mimo-9b-agent decodes at 6.57 tok/s (llama-bench tg128,
Vulkan, b11153) and has no recorded rate yet, so the stale value would size its
walls about a third too short. TestAmdGcnSeedsTheSeatItWasMeasuredOn pins the new
value (proven red at 10).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… generate)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dmmdea
dmmdea merged commit 4793b44 into main Sep 24, 2026
5 checks passed
@dmmdea
dmmdea deleted the feat/mimo-agent-seat-amd-gcn branch September 24, 2026 12:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant