Skip to content

fix(vllm): stop the Mooncake store from saving a load MultiConnector gave away - #282

Merged
thxCode merged 2 commits into
mainfrom
fix/vllm-mooncake-store-consumer-save
Sep 24, 2026
Merged

thxCode merged 2 commits into
mainfrom
fix/vllm-mooncake-store-consumer-save

Conversation

@thxCode

@thxCode thxCode commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Defect

On the published vllm0.29.0 runners, a decode engine configured as MultiConnector(MooncakeConnector, MooncakeStoreConnector) (both kv_consumer) exits on the first request whose prompt the prefill side has already written to the store:

AssertionError: Missing current block table for store request ...

Call chain: Scheduler.schedule -> _build_kv_connector_meta -> MultiConnector.build_connector_meta -> MooncakeStoreConnector.build_connector_meta -> MooncakeStoreScheduler._apply_current_save_block_ids.

Root cause

vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py at v0.29.0:

  1. The prompt is a hit for both children. MultiConnector.get_num_new_matched_tokens assigns the request to the first child with matched tokens (MooncakeConnector, async load) and update_state_after_alloc calls the store child with num_external_tokens=0, so the store's LoadSpec stays can_load=False.
  2. The request waits for remote KV and is scheduled neither as new nor as cached, so it lands in the "pending load specs not yet scheduled" branch (lines 366-391), which builds its ReqMeta with skip_save=None (line 388). With can_load=False, ReqMeta.from_request_tracker turns that into can_save=True.
  3. The core snapshots current block tables only for scheduled requests (vllm/v1/core/sched/scheduler.py, KVConnectorBlockState), so _apply_current_save_block_ids (line 424) finds no table and asserts.

The assertion was introduced by upstream vllm-project/vllm#51358 (6b110badbb), first released in v0.29.0. v0.27.1 has the same branch without the assertion.

Upstream has already fixed this in vllm-project/vllm#54643 (34b1e9f7a6), first released in v0.30.0 (also in v0.29.1rc0; no v0.29.x final carries it): the branch now skips a LoadSpec whose load this connector does not own (if load_spec is None or not load_spec.can_load: continue).

Fix

Backport the #54643 hunk verbatim, with line numbers rebased onto v0.29.0, by the two paths the MooncakeConnector Prometheus fix used (766ef2f, ecab586):

  • pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/: patches the published cuda13.0-vllm0.29.0, cuda12.9-vllm0.29.0 (amd64 and arm64) and rocm7.2-vllm0.29.0 (amd64) images in place. It is idempotent, asserts the patch landed, imports the patched module, and ends with the dependency probe and vllm-deps export stage, as the post operation README requires.
  • pack/{cuda,rocm}/patches/vllm/003_mooncake_store_rejected_pending_load.patch: carries the same hunk into new 0.29.0 builds, so that a rebuild does not ship the defect again and overwrite the in-place repair. Byte-identical between the two sides.

Why not skip_save=is_consumer, which the sibling branches use: it clears a consumer, but a kv_both engine still trips the same assertion. A request that has not computed anything has nothing to save under any role, and upstream's fix is what v0.30.0 ships.

Not affected: cann builds on vLLM 0.23.0, whose store scheduler has no such assertion, and vllm-ascend's kv pool runs its own scheduler instead of MooncakeStoreConnector's.

Interaction with #280

v0.30.0 still has skip_save=None in that branch (line 399), and the assertion is still there (line 444), but #54643's guard two lines above means the branch is only reached with can_load=True, where ReqMeta.from_request_tracker forces skip_save=True. It emits no save job, so the assertion cannot fire from this path; the reproduction below passes on an unpatched v0.30.0 tree. The patch is therefore meant for 0.29.0 only and does not apply to 0.30.0 (git apply --check: v0.29.0 forward passes; v0.30.0 forward fails, reverse passes).

#280 moves the pack path to 0.30.0, which already contains the fix. 003_*.patch is refused there (GNU patch: "Reversed (or previously applied) patch detected", exit 1), and the Dockerfile.vllm patch loop stops the build. Whichever of the two merges second needs to delete pack/{cuda,rocm}/patches/vllm/003_mooncake_store_rejected_pending_load.patch as part of the 0.30.0 bump. The post operation is unaffected: it pins 0.29.0 and skips a tree that already carries the fix.

Verification

  • git apply --check of 003_*.patch against v0.29.0 passes; against v0.30.0 the forward check fails and the reverse check passes, which shows the change is already upstream there.

  • GNU patch 2.7.6 applies 001, 002 and 003 in the build's order onto a v0.29.0 tree; the patched scheduler.py compiles.

  • Reproduction in the style of tests/v1/kv_connector/unit/test_mooncake_store_scheduler.py: a request with a rejected LoadSpec, parked unscheduled, with the block state the core would build (no entry for it).

    tree upstream #54643 regression test repro, kv_consumer repro, kv_both
    v0.29.0 fails AssertionError: Missing current block table ... AssertionError: Missing current block table ...
    v0.29.0 + 003 passes passes passes
    v0.29.0 + skip_save=is_consumer fails passes fails
    v0.30.0 passes passes passes

    The existing test_mooncake_store_scheduler.py suite (39 tests) passes on v0.29.0 + 003.

  • The post operation's script run over a v0.29.0 tree: the first run patches it (result byte-identical to the tree 003 produces), the second is a no-op. A v0.30.0 tree is left untouched. A stub scheduler.py is refused at the patch step.

  • expand_matrix.sh for the post operation yields exactly the five published platform tags listed in runner.py.json for vLLM 0.29.0.

  • pre-commit run on the changed files and pytest (228 passed).

Not verified here: building the post operation against the real published images. That runs in the Pack workflow with post_operation set, after this merges.

…gave away

A decode engine configured as MultiConnector(MooncakeConnector, MooncakeStoreConnector)
exits on the first request whose prompt the prefill side has already written to the
store:

    AssertionError: Missing current block table for store request ...

raised from MooncakeStoreScheduler._apply_current_save_block_ids, reached through
MultiConnector.build_connector_meta.

The prompt is a hit for both children. MultiConnector hands the request to the first
child with matched tokens, MooncakeConnector, which loads asynchronously, and calls the
store child with zero external tokens, so the store's LoadSpec stays can_load=False.
The request is then parked waiting for remote KV and is scheduled neither as new nor
as cached, which sends it into the store scheduler's "pending load specs not yet
scheduled" branch. That branch builds its ReqMeta with skip_save=None, so the rejected
LoadSpec becomes a save job; the core snapshots current block tables only for requests
it scheduled, and the assertion finds none for this one.

The assertion arrived with upstream #51358 and first shipped in 0.29.0; 0.27.1 has the
same branch without it. Upstream fixed the branch in #54643, first shipped in 0.30.0: it
skips a LoadSpec whose load this connector does not own. The post operation applies that
hunk verbatim, rebased onto the 0.29.0 line numbers. It is not skip_save=is_consumer,
which the sibling branches use: that would clear a consumer but leave a kv_both engine
tripping the same assertion, and a request that computed nothing has nothing to save
under any role.

Scope: the published 0.29.0 images, which are cuda13.0, cuda12.9 and rocm7.2. cann is
not affected: it builds on vLLM 0.23.0, whose store scheduler has no such assertion, and
vllm-ascend's kv pool runs its own scheduler rather than MooncakeStoreConnector's.

Verified against the vLLM v0.29.0 source tree. A reproduction in the style of
tests/v1/kv_connector/unit/test_mooncake_store_scheduler.py, which parks a request with
a rejected LoadSpec and hands the connector the block state the core would, raises the
assertion above for kv_consumer and kv_both before the change and passes after it, as
does upstream's own regression test from #54643; the existing store scheduler suite
still passes. The post operation's script was run over the tree twice, leaving it
byte-identical to the new-build patch the first time and untouched the second, left a
v0.30.0 tree untouched, and failed on a stub scheduler.py it has to refuse.

Signed-off-by: thxCode <thxcode0824@gmail.com>
The post operation beside this commit repairs the published 0.29.0 images. A rebuild
of 0.29.0 from the pack path would ship the defect again and overwrite that repair, so
the same hunk goes where new builds are assembled.

Both Dockerfile.vllm paths apply every patches/vllm/*.patch in order from the vLLM
site-packages directory and stop on the first that fails, so the fix is one more file on
each side, byte-identical between them as 001 and 002 already are. The hunk is upstream
#54643 verbatim, with line numbers rebased onto v0.29.0.

Verified with GNU patch 2.7.6 over a v0.29.0 tree in the order the build uses: 001, 002
and 003 all apply and the patched scheduler.py compiles. git apply --check accepts it
against v0.29.0 and refuses it against v0.30.0, where the reverse check passes because
upstream already carries the change.

That refusal is deliberate and loud: moving the pack path to 0.30.0 or later has to
delete this file, and a build that still carries it stops at the patch step rather than
shipping silently.

Signed-off-by: thxCode <thxcode0824@gmail.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a patch for vLLM 0.29.0 to resolve an AssertionError occurring in the MooncakeStoreScheduler when a LoadSpec is rejected by MultiConnector. The changes include updated Dockerfiles for CUDA and ROCm, new patch files, and updates to the project's matrix and documentation. I have no feedback to provide.

@thxCode
thxCode merged commit 80034d1 into main Sep 24, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant