Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
45 commits
Select commit Hold shift + click to select a range
c0cbae5
feat(BACKEND-ROCM): register a ROCm attention backend for kROCM
tbrasser Aug 16, 2026
b493f49
fix(ENG-EXPERT-STREAM): the grouped CUDA dispatch launched nothing an…
localai-bot Aug 16, 2026
ef66692
fix(#1055): fold the two Darwin paragraphs into rows and take main ou…
localai-bot Aug 16, 2026
ec0e410
perf(LTX25-DECODE-THREADS): the video VAE decode had 20 cores and use…
localai-bot Aug 16, 2026
591d24f
perf(MODEL-MUSIC-MUSIC3): the CPU arm gets its cores back — three hos…
localai-bot Aug 16, 2026
0b0b890
feat(LTX25-T2A-ONE-STAGE): text-to-audio, the one-stream forward, and…
localai-bot Aug 16, 2026
d607fec
fix(VT-FP8-QUANT-ARCH-GATE): register QuantFp8Static from an uncondit…
localai-bot Aug 16, 2026
1000264
perf(MODEL-MUSIC-MUSIC3): the vocoder convolution chain is 10.7x, and…
localai-bot Aug 16, 2026
a332fb9
fix(ENG-EXPERT-STREAM): the step clock had no caller, so the decode n…
localai-bot Aug 16, 2026
cf4e672
fix(ltx2): cast positions double/float to avoid MSVC C4244 (#968)
harleywilsoneng Aug 16, 2026
e9dfa63
record(ORACLE-LLAMACPP-OWED-CPU): the boundary, the decision, and thr…
localai-bot Aug 17, 2026
9143196
docs(FEATURES): the registry has 40 architectures, and the table said 38
localai-bot Aug 17, 2026
b5756ea
record(LTX25-RESOLUTION-ENVELOPE): 704x448/25f completes, and the siz…
localai-bot Aug 17, 2026
0bac476
fix(MODEL-MUSIC3): the italic unwrap CONSUMED its trailing neighbour,…
localai-bot Aug 17, 2026
6abc769
feat(#810 A2-Q2a): NemotronH's 23 MoE blocks reach the device on an N…
localai-bot Aug 17, 2026
22056e2
fix(FIX-NAS-PATH-1073): resolve every live checkpoint default through…
localai-bot Aug 17, 2026
e5fedf2
spec(#810 A2-P): the paged forward, and the two things the governing …
localai-bot Aug 17, 2026
281e6a1
record(ROAD-V1-LTX25): four upstream pipelines nobody had filed, and …
localai-bot Aug 17, 2026
8fa405b
fix(#595): doc-checkpoint asks whether the registry moved, not whethe…
localai-bot Aug 17, 2026
9b3317c
fix(qwen3.5): the capture #1054 dropped was not redundant, and MSVC s…
localai-bot Aug 17, 2026
268da6b
fix(ENG-EXPERT-STREAM): the liveness line could not print the zero th…
localai-bot Aug 17, 2026
daeff67
feat(LTX25-GUIDED-VIDEO): the guided video denoiser, and the sentence…
localai-bot Aug 17, 2026
a6df727
feat(#810 A2-P): NemotronH decodes on the runner's own pages, so step…
localai-bot Aug 17, 2026
2e02524
docs(FEATURES): count all 12 parsers (#1103)
localai-org-maint-bot Aug 17, 2026
994d30b
gate(SPEC-MTP-K-GT-1-DGX): the padded control is paid on real weights…
localai-bot Aug 17, 2026
589abad
feat(#672): vt::Conv1d and vt::ConvTranspose1d, with a CUDA provider …
localai-bot Aug 17, 2026
4d77486
feat(LTX25-RES2S-LOOP): the res_2s loop, its second evaluation, and t…
localai-bot Aug 17, 2026
559973c
fix(rocm/gemma4): #837 GetBlas dual-slot TLS + host lifetime seam
bakon11 Aug 17, 2026
8460764
feat(MODEL-MUSIC-MUSIC3): the 2.4B fp32 DiT onto the device, staged o…
localai-bot Aug 17, 2026
d1e5e9b
feat(LTX25-A2VID-RECIPE): the audio-to-video recipe, and the phase fi…
localai-bot Aug 17, 2026
ef7f6f6
policy(ENV-GPU-LEASE-METHODOLOGY): a GPU lease is the rule, and the c…
localai-bot Aug 17, 2026
a7583ac
perf(rocm): head_dim=128 decode arm -- the ROCm half of #382, default…
joral Aug 17, 2026
ac50579
record(QUANT-GGUF-KEEPQ-LOADER): VT_GGUF_KEEP_F16 stays default ON, s…
localai-bot Aug 17, 2026
04b58bf
fix(GATE-ISSUE-INDEX-TABLE-SHAPE): count the index's cells, and repai…
localai-bot Aug 17, 2026
f6f1af0
record(ENV-LEASE-RUNTIME-STAGING): a relocated CUDA runtime starts in…
localai-bot Aug 17, 2026
4ae0f54
feat(LTX25-PHASE-LORA): the adapter set belongs to a phase, and the D…
localai-bot Aug 17, 2026
c83b969
feat(#810 A3): NemotronH gets its ABI driver, and the reason its gate…
localai-bot Aug 17, 2026
3604c0e
fix(BACKEND-ROCM): re-derive upstream anchors at 555967922 and mirror…
tbrasser Aug 17, 2026
76f2a6d
fix(ENG-EXPERT-STREAM): refuse a GGUF the device cannot hold, at load…
localai-bot Aug 17, 2026
40a796a
feat(LTX25-BF16-DIT): the DiT that most LTX-2.5 pipelines need, and t…
localai-bot Aug 17, 2026
4fe0f2f
fix(ENG-RELEASE-WINDOWS): the main baseline runs the two MSVC gates i…
localai-bot Aug 17, 2026
79ff8f3
feat(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): capture the Qwen3 decode…
lu-zero Aug 17, 2026
affc2a7
feat(LTX25-TI2VID-RECIPE): the plain two-stage pipeline, on the sched…
localai-bot Aug 17, 2026
eceb666
merge: origin/main into row/BACKEND-ROCM-ATTN-REGISTER
mudler Aug 17, 2026
935872d
docs(rocm): trim the FEATURES cell under the cap, and say what ROCM_A…
mudler Aug 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions .agents/NOW.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# NOW — the one-Read resume surface

<!-- now-updated: 2026-08-11 -->
<!-- now-updated: 2026-08-17 -->

Snapshot, not log. History is git; evidence:
[parity ledger](parity-ledger.md), and benchmarks. Budget: 100 lines / 6,000
Expand Down Expand Up @@ -65,8 +65,10 @@ devices IN SCOPE (`ROAD-V1-D6`).
- Mirror vLLM; never ask how a feature should behave.
- `nsys` BOTH sides, SAME tool, before any perf claim; cross-tool comparisons
never establish invocation parity; whole-run sums mix prefill.
- GPU: park `local-ai-worker`, flock `$HOME/gpu.lock`, single-load
steady-state, never reload per rep, named tmux.
- GPU: claim a lease with `rc run` or `rc hold`. Never `ssh` to a box, because
that makes the fleet report it free while you are on it. The flock now lives
INSIDE the lease ([environment](environment.md)). Single-load steady-state,
never reload per rep, named tmux.
- Never weaken a checker to pass; repair the record.
- Work happens in its own worktree on a task branch; the shared checkout stays
clean on `main`, never a work surface. Land via `row/*` PR or authorized
Expand Down
5 changes: 3 additions & 2 deletions .agents/backend-matrix.md

Large diffs are not rendered by default.

587 changes: 587 additions & 0 deletions .agents/benchmark-record.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/claims/CLAIM-ROCM-DECODE-ATTN-D128.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,4 @@

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-ROCM-DECODE-ATTN-D128` | `BACKEND-ROCM` (`ACTIVE`) | Claude Code (sonnet-5), helper role | worktree `rdna3-kernel-porting-b9ec47`, real gfx1200 hardware (AMD Radeon RX 9060 XT, RDNA4, 32 CU), `$GPU_LOCK` respected | `row/ROCM-DECODE-ATTN-D128-SPEC` (this spec; the implementation follows on `row/ROCM-DECODE-ATTN-D128-IMPL`, stacked), base `main` `fafa16f0`; issue [#382](https://github.com/mudler/vllm.cpp/issues/382) (the ROCm half; the CUDA half landed as [PR #425](https://github.com/mudler/vllm.cpp/pull/425), `66399617`), motivated by [#488](https://github.com/mudler/vllm.cpp/issues/488). NOTE: #382 is filed against the cross-backend kernel row (state `ANCHOR-BACKFILL`), while this claim's Row ID is the `ACTIVE` backend row whose code it edits — `check-agent-record` requires an active claim to name a `SPIKE`/`ACTIVE` row, so the two deliberately differ | Owns ONLY: the `LoadRowEplBf16`/`StoreRowEplBf16` `EPL=4` case, the `VT_ATTN_DECODE_D128` gate (default OFF, same flag/default/reason as the merged CUDA arm), the `bf16_decode_opt`/`decode_gqa` gate extensions and the two `d==128` launch-dispatch branches in `src/vt/rocm/rocm_paged_attn.hip`; the new "Qwen3 geometry (bf16, GQA 2, head_dim 128)" case in `tests/vt/test_backend_cross_device.cpp` and its two flag-on ctest registrations in `tests/CMakeLists.txt`; `.agents/specs/rocm-decode-attn-d128.md` and this claim file. **NON-COLLISION:** disjoint from `CLAIM-ROCM-SKINNY-GEMM-GFX1200` (different files: `rocm_skinny_gemm.hip`/`rocm_matmul_hipblaslt.hip` vs `rocm_paged_attn.hip`), not stacked on any other branch. EXCLUDED: **the flip to default-ON on either backend** (owes the near-tie razor + distributional gate + golden regen, and per the spec §5 cross-arch reversal must be argued per backend — this is what keeps #382 open), rocWMMA for `d=128` (separate claim, separate spec, separate issue), `qg=4`/`qg=8` GQA fusion at any `d` (pre-existing, board-independent gap), any `d=128` prefill path, and the 8 pre-existing unrelated `ctest` failures (`vt: no kernel for op 63 on device type 5`) | `ACTIVE` | 2026-08-12 — **reconciled against the existing record before landing**, per the re-verify-before-claiming rule: #382 already named this exact defect and PR #425 had already merged the CUDA half, so this became a mirror of merged work rather than new design, and was re-gated from default-ON to **default OFF behind `VT_ATTN_DECODE_D128`** — the merged arm's own flag, default and stated reason (warp-strided online softmax reduces the KV sequence in a different order, so a greedy anchor can move at an exact bf16 tie; OFF keeps every golden byte-identical). gfx1200-verified: `ctest -R 'rocm\|cross_device'` **6/6** including two new flag-on registrations (verified non-vacuous: 1 case, 6 assertions, not zero); full `ctest` 385/393 with the 8 failures independently confirmed pre-existing. Gate exercised **both directions on one binary** — Qwen3-0.6B @1024 ctx TPOT 44.82/44.82 ms OFF vs 12.80/12.60 ms ON = **3.53x**; decode throughput +42.7% / +25.0% / +17.8% on 0.6B / 1.7B / 4B. **Carried finding:** #382 measured this same `EPL=4` arm **1.6x slower** on sm_110 where we measure it 3.5x faster — recorded, not reconciled; it is why the default-ON flip must be argued per backend. Rebased from `bbc482a2` onto `main` `fafa16f0` (167 commits), which required reformatting `Assisted-by` for the `check-commit-trailers` gate that landed in between, and de-linking §7's forward reference to the rocWMMA spec — that spec now lands on its own branch, so a markdown link to it fails `check-agent-record` as a dangling link. Spec content otherwise byte-identical. Re-gated on the new base, gfx1200: build 783/783, `ctest -R 'rocm\|cross_device'` 6/6, the new case non-vacuous under both flags (1 case, 6 assertions), full `ctest` with 8 pre-existing `kSharedExpertGate` (`OpId(63)`) failures owed to unmerged PR #509. `agent-preflight` fails 11, set-identical to a clean `fafa16f0` baseline. Spec PR open; implementation PR follows. |
| `CLAIM-ROCM-DECODE-ATTN-D128` | `BACKEND-ROCM` (`ACTIVE`) | Claude Code (sonnet-5), helper role | worktree `rdna3-kernel-porting-b9ec47`, real gfx1200 hardware (AMD Radeon RX 9060 XT, RDNA4, 32 CU), `$GPU_LOCK` respected | `row/ROCM-DECODE-ATTN-D128-IMPL` (the implementation; its spec landed from `row/ROCM-DECODE-ATTN-D128-SPEC` as [PR #564](https://github.com/mudler/vllm.cpp/pull/564), squashed to `373aa125`), rebased off the now-merged spec commits onto `main` `2784dd7b`; issue [#382](https://github.com/mudler/vllm.cpp/issues/382) (the ROCm half; the CUDA half landed as [PR #425](https://github.com/mudler/vllm.cpp/pull/425), `66399617`), motivated by [#488](https://github.com/mudler/vllm.cpp/issues/488). NOTE: #382 is filed against the cross-backend kernel row (state `ANCHOR-BACKFILL`), while this claim's Row ID is the `ACTIVE` backend row whose code it edits — `check-agent-record` requires an active claim to name a `SPIKE`/`ACTIVE` row, so the two deliberately differ | Owns ONLY: the `LoadRowEplBf16`/`StoreRowEplBf16` `EPL=4` case, the `VT_ATTN_DECODE_D128` gate (default OFF, same flag/default/reason as the merged CUDA arm), the `bf16_decode_opt`/`decode_gqa` gate extensions and the two `d==128` launch-dispatch branches in `src/vt/rocm/rocm_paged_attn.hip`; the new "Qwen3 geometry (bf16, GQA 2, head_dim 128)" case in `tests/vt/test_backend_cross_device.cpp` and its `VT_ATTN_DECODE_D128` flag-on ctest registration in `tests/CMakeLists.txt` (the second, `VT_ATTN_DECODE_WMMA`, moved to the rocWMMA branch with the arm it gates); `.agents/specs/rocm-decode-attn-d128.md` and this claim file. **NON-COLLISION:** disjoint from `CLAIM-ROCM-SKINNY-GEMM-GFX1200` (different files: `rocm_skinny_gemm.hip`/`rocm_matmul_hipblaslt.hip` vs `rocm_paged_attn.hip`), not stacked on any other branch. EXCLUDED: **the flip to default-ON on either backend** (owes the near-tie razor + distributional gate + golden regen, and per the spec §5 cross-arch reversal must be argued per backend — this is what keeps #382 open), rocWMMA for `d=128` (separate claim, separate spec, separate issue), `qg=4`/`qg=8` GQA fusion at any `d` (pre-existing, board-independent gap), any `d=128` prefill path, and the pre-existing unrelated `ctest` failures (`vt: no kernel for op SharedExpertGate` on ROCm, plus a missing `shellcheck`, an mmap-RSS assertion and a JSON type error) | `ACTIVE` | 2026-08-12 — **reconciled against the existing record before landing**, per the re-verify-before-claiming rule: #382 already named this exact defect and PR #425 had already merged the CUDA half, so this became a mirror of merged work rather than new design, and was re-gated from default-ON to **default OFF behind `VT_ATTN_DECODE_D128`** — the merged arm's own flag, default and stated reason (warp-strided online softmax reduces the KV sequence in a different order, so a greedy anchor can move at an exact bf16 tie; OFF keeps every golden byte-identical). gfx1200-verified: `ctest -R 'rocm\|cross_device'` **6/6** including two new flag-on registrations (verified non-vacuous: 1 case, 6 assertions, not zero); full `ctest` 385/393 with the 8 failures independently confirmed pre-existing. Gate exercised **both directions on one binary** — Qwen3-0.6B @1024 ctx TPOT 44.82/44.82 ms OFF vs 12.80/12.60 ms ON = **3.53x**; decode throughput +42.7% / +25.0% / +17.8% on 0.6B / 1.7B / 4B. **Carried finding:** #382 measured this same `EPL=4` arm **1.6x slower** on sm_110 where we measure it 3.5x faster — recorded, not reconciled; it is why the default-ON flip must be argued per backend. Rebased from `bbc482a2` onto `main` `fafa16f0` (167 commits), which required reformatting `Assisted-by` for the `check-commit-trailers` gate that landed in between, and de-linking §7's forward reference to the rocWMMA spec — that spec now lands on its own branch, so a markdown link to it fails `check-agent-record` as a dangling link. Spec content otherwise byte-identical. Re-gated on the new base, gfx1200: build 783/783, `ctest -R 'rocm\|cross_device'` 6/6, the new case non-vacuous under both flags (1 case, 6 assertions), full `ctest` with 8 pre-existing `kSharedExpertGate` (`OpId(63)`) failures owed to unmerged PR #509. `agent-preflight` fails 11, set-identical to a clean `fafa16f0` baseline. **2026-08-14 — spec LANDED as PR #564 (`373aa125`); this claim now tracks the implementation.** Rebased off the two now-squashed spec commits onto `main` `2784dd7b`; the commit is source-only (3 files) and carries no forward reference to the rocWMMA flag, so it stands alone. Re-gated on that base, gfx1200, `$GPU_LOCK` held: build 1220/1220; `ctest -R 'rocm\|cross_device'` **5/5** (5 not 6 — the `VT_ATTN_DECODE_WMMA` registration left with its arm); flag A/B on ONE binary re-measured **3.47x** (Qwen3-0.6B @1024 ctx, 45.15/45.13 ms OFF vs 12.97/13.04 ms ON), holding the 3.53x from the old base across 76 commits of drift. Full `ctest` 448/455 with **7** failures, and those 7 are now PROVEN pre-existing rather than argued: a clean `main` `2784dd7b` worktree, built from source with none of this code, fails the identical set (only `test_op_parity`'s index shifts 403→404, from the added registration). `agent-preflight` fails 9, a strict SUBSET of that same baseline's 10 (differing only by `role-undeclared`). Note `origin/main` (the `joral` fork) is 75 commits behind `upstream/main`, so preflight's range gates grade 76 commits of which 75 are other people's — `check-commit-trailers` and `check-doc-checkpoint` both pass against `upstream/main`, the base the spec actually merged to. **Fresh evidence 2026-08-14, SUPERSEDING the `+42.7% / +25.0% / +17.8%` figures above** — those came from a 128-token-context stash-based A/B; every number here is 1024-token synthetic prompt, 128 generated, greedy, seed 0, one binary, `$GPU_LOCK` held, 2 reps per cell agreeing within ~1%. Four-model TPOT OFF→ON: Qwen3-0.6B 42.53→11.78 ms (**3.61x**, `qg=2` fused), Qwen3-1.7B 52.85→21.93 ms (**2.41x**, `qg=2` fused), Qwen3-4B 81.89→39.22 ms (**2.09x**, `qg=4` per-head — no GQA fusion at any `d`, so this isolates the `EPL` widening from the fusion), Qwen3.5-0.8B 23.76→23.55 ms (**1.01x**). The last is the **NEGATIVE CONTROL** and it earned its keep: its `head_dim` is 256, so the `d == 128` gate provably cannot reach it, yet its first OFF rep landed a 33% outlier at 31.14 ms — a blind 2-rep average would have reported a bogus ~1.2x "win" for a model the flag cannot affect. Re-run 3x it gives 23.86/23.75/23.68 against ON's 23.52/23.57. End-to-end output throughput rises less than TPOT on the same runs (0.6B 2.48x, 1.7B 2.05x, 4B 2.02x) because these carry a 1024-token prefill the flag does not touch; TPOT isolates decode, throughput dilutes it. **Qwen3-1.7B concurrency sweep** (`--num-prompts` = 2x concurrency), throughput tok/s OFF→ON (ratio): c1 12.89→24.66 (1.91x), c2 23.27→47.45 (2.04x), c4 39.10→86.86 (2.22x), c8 58.97→147.35 (**2.50x**), c16 78.43→227.08 (**2.90x**); TPOT ratio over the same points 2.40x→3.18x. **The advantage GROWS with concurrency rather than compressing** — the opposite of the prediction made before the run, which reasoned that a tiny c1 grid flatters the fast kernel. The dominant effect is the reverse: from c8 to c16 the fallback scales only 1.33x against the arm's 1.54x, and scaling efficiency at c16 relative to perfect-linear-from-c1 is **38% OFF against 58% ON**. `PagedAttnOnline` is therefore the batch-scaling bottleneck, not merely slow per call, and the win is largest in the regime a server actually runs in. The c1 row reproduces the independent four-model sweep to within ~1% (52.85/21.93 vs 53.40/22.26), a cross-check on run-to-run stability. **Caveats:** single board; `--input-len` builds synthetic tokens, so all of the above is a decode-path A/B and not a serving benchmark. Implementation PR not yet opened. |
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-ROCM-GEMMA4-GETBLAS-DUALSLOT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-ROCM-GEMMA4-GETBLAS-DUALSLOT

| Claim | Row IDs | Agent | Worktree | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-ROCM-GEMMA4-GETBLAS-DUALSLOT` | `BACKEND-ROCM` (slug `ROCM-GEMMA4-GETBLAS-DUALSLOT`, issue #837) | hermes-vllm (lab), helper | `/home/don/llms/vllm.cpp-getblas` | `row/ROCM-GEMMA4-GETBLAS-DUALSLOT` | Owns ONLY: `GetBlas` `tls_slots[2]` in `src/vt/rocm/rocm_matmul_hipblaslt.hip` plus host lifetime seam tests. **EXCLUDED:** Launch/Finish (#839), indexed T (#838), #697 / `rocm_paged_attn.hip`. Independent history from the abandoned combined branch `row/ROCM-GEMMA4-XDEV-MOE`. | `IMPLEMENTING` | 2026-08-15 — 6195 production capture hook load-bearing |
Loading
Loading