update(speculative): gate CUDA prompt-lookup drafting on shadow probes - #2092
Merged
Merged
Conversation
On GB10 prompt lookup ran prose 3-21% below plain decoding (#2091). A synchronous verify forward there costs, in pipelined steps, 1.14-1.35 at width 2, 1.41-1.62 at width 3, and 2.5-3.6 from width 5 up (widths 5-7 cost more than 8, where `qmm_sm80` takes over), and every drafted round also drained the pipeline. The PR #2074 governor kept probing prose with such blocks every 4-32 rounds. `DraftPolicy::Gated`, the default on CUDA builds, uses only a narrow block (2 proposals) or the full one, starts narrow and on probation, and never ends a pause on a timer: while paused it keeps looking up and resumes only when two proposed tokens in a row come true unverified (`ShadowProbe`), costing a hash probe instead of a verify forward. `DraftPolicy::Graded` keeps PR #2074's rules unchanged and stays the default elsewhere, since they were tuned on Apple Silicon and no Metal measurement backs a change. `--prompt-lookup-policy auto|graded|gated` selects either for A/B, and the summary line reports `shadow_confirmations`. `examples/verify_width_cost.rs` measures the per-width cost table quoted in `DraftPolicy`. Tests: gated start, probation, timer-free pause, narrow/full widths, shadow settling, policy defaults, CLI parsing; the rollback parity matrix now runs both policies and requires a shadow-confirmed resume. Five deliberate mutations of the gated logic each fail them. Refs #2091
Review fixes for #2092: - A paused `DraftPolicy::Gated` governor now pipelines its plain rounds at once instead of after two synchronous ones: it cannot propose until a shadow probe settles, which takes at least two emitted tokens, so pipelining delays nothing. - Shadow lookups ask for exactly the two tokens a probe settles on. - Rollback parity totals are kept per policy, and Graded must never report a shadow confirmation. - New exact test: a script model copies its prompt, breaks off for one token and resumes; the gated loop must confirm the resumed copy exactly once, pipelined and on the synchronous history-sampler path. A probe recorded one position early or late now fails it. - `DraftPolicy` is `#[non_exhaustive]`, its doc separates the budget from the verify width and states the kernel boundaries once, the end-of-decode trace reports `shadow_confirmations`, and `verify_width_cost` refuses models whose caches cannot be trimmed. Refs #2091
The first GB10 matrix left Qwen3-8B's email at 0.96x of plain decoding with 6 drafted rounds, 3 landed tokens and no pause: rounds that landed one stray token reset the miss count and lifted probation, so the misses between them were never caught. Probation now ends only when a round lands a whole narrow block (2 proposals); one landed token still resets the miss count but leaves probation on, so the next miss pauses. Refs #2091
The second GB10 matrix left Qwen3-8B's email at 0.97x. A round-by-round trace showed why: the reply restates a short fragment of the prompt, one narrow block landed whole, the governor switched to the full block, and the next two full blocks (width 8, about 2.5 pipelined steps on 8B) landed nothing as the fragment ended. `GATED_FULL_AT` rises from 1.5 to 1.75, so two narrow blocks in a row must land whole (four confirmed tokens) before a full block is spent. Edits keep full blocks once they land several tokens, since the average then stays above the threshold through a miss. Refs #2091
Reverts 7224525. On the third GB10 matrix the higher `GATED_FULL_AT` left Qwen3-8B's email where it was (0.97x, graded 0.97x) and changed the verify schedule enough that the Qwen3-1.7B and Qwen3-4B emails, byte-identical to plain decoding at 1.5 and under Graded, diverged from it (near-tie flips in the multi-token verify). No measured gain, a lost parity criterion: back to 1.5. Refs #2091
Adds `docs/benchmark_results/prompt-lookup-governor-gb10-2026-10-02.md` with the interleaved harness, prompts, raw results of all three matrices, the load log, and the per-width verify cost on four models, and links it from `docs/benchmarks.md`. `DraftPolicy`'s doc now quotes the committed release-build ratios instead of the earlier test-fast probe. Final matrix (`34c627d7`, greedy, median of 3, plain/graded/gated interleaved): story 0.98x to 1.00x on all four models (graded 0.83x to 0.93x); write 0.98x to 1.02x except Qwen3-8B at 0.97x (graded 0.98x in the same run); edit and summary gains kept within 4% of graded or raised; every reply graded keeps byte-identical to plain decoding stays identical under gated. Refs #2091
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Prompt lookup slowed prose decoding 3 to 21% on GB10 (#2091). This adds
DraftPolicy::Gated, the new default on CUDA builds, and keeps PR #2074's governor unchanged asDraftPolicy::Gradedfor Metal and ROCm. On GB10, stories go from 0.83x-0.93x to 0.98x-1.00x of plain decoding, edit and summary gains are kept or raised, and greedy parity is preserved. One target is not met: Qwen3-8B's email row stays at 0.97x (graded 0.98x in the same run).Why
A synchronous verify forward on GB10 (affine 4-bit, MLX pin
81ba1c6a, release build) costs, in pipelined steps for Qwen3-1.7B / Qwen3-8B: width 2 at 1.37 / 1.15, width 3 at 1.76 / 1.59, width 8 at 3.91 / 3.30, with widths 5 to 7 close to or above width 8 (qmv's 8-row multirow instantiation, thenqmm_sm80from 8 rows). A drafted round also drains the pipeline. Graded, tuned on an M4 Pro, re-probed prose with blocks up to width 8 every 4 to 32 rounds.What changed
DraftPolicy::Gated: blocks are narrow (2 proposals) or full (max_draft), widening after a narrow block lands whole. A reply starts narrow and on probation, and probation ends only on a round that lands a whole narrow block. A pause has no timer: while paused the loop keeps looking up and checks each proposal against the tokens decoding emits next (ShadowProbe); drafting resumes, on probation, only after two proposed tokens in a row came true. Paused rounds pipeline at once.DraftPolicy::Gradedmakes exactly PR feat(speculative): add prompt-lookup decoding to mlxcel generate #2074's decisions and stays the default off CUDA: no Apple Silicon machine was reachable, so Metal keeps the rules it was tuned with.--prompt-lookup-policy auto|graded|gatedfor A/B on any backend;shadow_confirmationson the[Prompt lookup]line and in the decode trace.examples/verify_width_cost.rsmeasures the per-width costs above. The verify path and the in-flight transition are unchanged.Results (GB10)
Final matrix on
34c627d7(the tree this PR ships), greedy, median of 3, plain/graded/gated interleaved; full tables, counters, parity and method indocs/benchmark_results/prompt-lookup-governor-gb10-2026-10-02.md.Every reply graded keeps byte-identical to plain decoding stays identical under gated. A third matrix with a stricter full-block threshold left Qwen3-8B's email at 0.97x and broke parity on two email rows, so it was reverted (
d678c940); the record explains the remaining cost and the round-loop change that would address it.Changes during review
No CRITICAL or HIGH findings. Applied: immediate pipelining while paused, an exact probe-alignment test, per-policy parity totals, two-token shadow lookups, budget-vs-width wording, and a trim guard in the width probe.
Validation
speculative::prompt_lookup: 39 passed; deliberate mutations of the gated rules, probe alignment and probation each fail them.mlxcelbin: 96 passed;dead_doc_pointerspassed; clippy-D warnings(lib, tests, bins, the new example) and fmt clean.Refs #2091