fix(server): close residual paged-budget gaps left by #2077 - #2095
Merged
Merged
Conversation
Per-layer accounting: a freed paged block stays on its layer's free list and acquire reuses only that list before minting under the pool-wide cap, so `budget - live` over-reported once the lists diverged and a layer could run dry while the figure was positive. `free_block_budget` now reports the largest uniform per-layer step every layer can acquire (own free blocks plus a share of the mint headroom), capped at `budget - live`. Admission, decode reservation and the lookahead prime all read it. Chunk reservation: `continue_chunked_prefill` reserves its chunk's blocks (tile padding included) against the pool figure before the forward. The set-aside figure saturates at 0, so once the pool fell below a parked prefill's reservation nothing restored it and the chunk could hit an exhausted pool. The reservation evicts cold prefixes and then preempts with floor 0, and covers both the MixedStep tick and the #1011 grant. The reservation itself now counts the next chunk's padding. Admission watermark: `--kv-admission-watermark` (env `MLXCEL_KV_ADMISSION_WATERMARK`) keeps a fraction of the block budget free at admission while a decode batch is live, evicting cold prefixes toward it before deferring. The batched window and the single-sequence gate charge against the same figure. The GB10 pressure sweep at higher watermarks exposed two #2077 accounting gaps, fixed here: admission charged an adopted request for its whole prompt although the adopted blocks were already live (a prefix over half the budget wedged the queue with an empty batch), and blocks pinned by queued adoptions were unreachable by reclaim (a lone row was shed). Admission now charges the suffix, and decode and chunk reclaim drop a queued request's adoption after cold prefixes and before preempting. Refs #2088
GB10 sweep with meta-llama-3.1-8b-instruct-4bit, a 4444-block budget, max batch 4, and three concurrent 3-turn conversations of 600 tokens, each value run twice: 0 gave 10 preemptions, 0.01 and 0.02 gave 6, 0.05 gave 5, and 0.10 and 0.15 gave 1, with all 9 turns completed in every run. 0.01 is the smallest grid value that lowers the count, so it becomes the default. Documents `--kv-admission-watermark` and `MLXCEL_KV_ADMISSION_WATERMARK` in CONTINUOUS_BATCHING.md and environment-variables.md, and adds the EN/KO technical report. Refs #2088
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes the three paged-budget gaps #2077 left, and two #2077 accounting gaps that the watermark measurement exposed.
budget - liveover-reported once the lists diverged.free_block_budgetnow reports the largest uniform per-layer step every layer can acquire (own free blocks plus a share of the mint headroom), capped atbudget - live. Unchanged for balanced lists. Re-tagging blocks across layers was rejected because pool rows are per-layer slabs.continue_chunked_prefillreserves its chunk's blocks (padding included) against the pool figure before the forward, evicting cold prefixes and then preempting with floor 0. The set-aside figure saturates at 0, so nothing restored a reservation the pool had fallen below. Covers theMLXCEL_MIXED_STEPtick and the fix(server): a chunked prefill is starved until the decode batch drains #1011 grant.--kv-admission-watermark F/MLXCEL_KV_ADMISSION_WATERMARK(0.0 to 0.5, default 0.01) keepsF * budgetblocks free at admission while rows decode, reclaiming toward it before deferring. Off with an empty batch, so nothing that fits the budget is refused.Measurement (GB10)
meta-llama-3.1-8b-instruct-4bit,
--kv-cache-budget 600000000(4444 blocks), paged, max batch 4, three concurrent 3-turn conversations of 600 tokens. Each branch value ran twice with identical counts; 0.10 and 0.15 ran once.0.01 is the smallest grid value that lowered the count.
MLXCEL_MIXED_STEP=1at 0.01: 21 mixed steps, 9/9 turns, no panics. Default launch (auto budget): no reclaim, prompt-cache hits intact.Test plan
server::batch,server::prompt_cache,runtime_settings,cli_input,config_tests,commands::serve(822 passed); mlxcel-corepaged,budget(282 passed)-D warnings(root and-p mlxcel-core) and fmt cleangpu-lockNot validated: the chunk reservation's reclaim path never fired on the real server (unit tests only); the Metal M5+ padding path; the queued-adoption drop fired twice end to end. Details in
TECHNICAL_REPORTS/2088-paged-budget-gaps-20261002.en.md.Closes #2088