From b1126f743b7c6fa392b70abf47e7ca711e8e4383 Mon Sep 17 00:00:00 2001 From: thxCode Date: Thu, 24 Sep 2026 08:54:31 +0800 Subject: [PATCH 1/2] fix(vllm): stop the Mooncake store from saving a load MultiConnector gave away A decode engine configured as MultiConnector(MooncakeConnector, MooncakeStoreConnector) exits on the first request whose prompt the prefill side has already written to the store: AssertionError: Missing current block table for store request ... raised from MooncakeStoreScheduler._apply_current_save_block_ids, reached through MultiConnector.build_connector_meta. The prompt is a hit for both children. MultiConnector hands the request to the first child with matched tokens, MooncakeConnector, which loads asynchronously, and calls the store child with zero external tokens, so the store's LoadSpec stays can_load=False. The request is then parked waiting for remote KV and is scheduled neither as new nor as cached, which sends it into the store scheduler's "pending load specs not yet scheduled" branch. That branch builds its ReqMeta with skip_save=None, so the rejected LoadSpec becomes a save job; the core snapshots current block tables only for requests it scheduled, and the assertion finds none for this one. The assertion arrived with upstream #51358 and first shipped in 0.29.0; 0.27.1 has the same branch without it. Upstream fixed the branch in #54643, first shipped in 0.30.0: it skips a LoadSpec whose load this connector does not own. The post operation applies that hunk verbatim, rebased onto the 0.29.0 line numbers. It is not skip_save=is_consumer, which the sibling branches use: that would clear a consumer but leave a kv_both engine tripping the same assertion, and a request that computed nothing has nothing to save under any role. Scope: the published 0.29.0 images, which are cuda13.0, cuda12.9 and rocm7.2. cann is not affected: it builds on vLLM 0.23.0, whose store scheduler has no such assertion, and vllm-ascend's kv pool runs its own scheduler rather than MooncakeStoreConnector's. Verified against the vLLM v0.29.0 source tree. A reproduction in the style of tests/v1/kv_connector/unit/test_mooncake_store_scheduler.py, which parks a request with a rejected LoadSpec and hands the connector the block state the core would, raises the assertion above for kv_consumer and kv_both before the change and passes after it, as does upstream's own regression test from #54643; the existing store scheduler suite still passes. The post operation's script was run over the tree twice, leaving it byte-identical to the new-build patch the first time and untouched the second, left a v0.30.0 tree untouched, and failed on a stub scheduler.py it has to refuse. Signed-off-by: thxCode --- .../cuda/Dockerfile | 101 ++++++++++++++++++ .../matrix.yaml | 42 ++++++++ .../rocm/Dockerfile | 101 ++++++++++++++++++ pack/.post_operation/README.md | 1 + 4 files changed, 245 insertions(+) create mode 100644 pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/cuda/Dockerfile create mode 100644 pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/matrix.yaml create mode 100644 pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/rocm/Dockerfile diff --git a/pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/cuda/Dockerfile b/pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/cuda/Dockerfile new file mode 100644 index 0000000..c4cd5ba --- /dev/null +++ b/pack/.post_operation/20260924_vllm_patch_mooncake_store_pending_load/cuda/Dockerfile @@ -0,0 +1,101 @@ +ARG CMAKE_MAX_JOBS +ARG CUDA_VERSION=12.9 +ARG VLLM_VERSION=0.29.0 + +FROM gpustack/runner:cuda${CUDA_VERSION}-vllm${VLLM_VERSION} AS vllm +SHELL ["/bin/bash", "-eo", "pipefail", "-c"] + +ARG TARGETPLATFORM +ARG TARGETOS +ARG TARGETARCH + +## Stop MooncakeStoreConnector from saving a load MultiConnector rejected +## +## With MultiConnector(MooncakeConnector, MooncakeStoreConnector) on the decode side, a prompt the +## prefill side has already written to the store is a hit for BOTH children. MultiConnector gives +## the request to MooncakeConnector, which loads asynchronously, and calls the store child with +## zero external tokens, so the store's LoadSpec stays can_load=False. The request is then parked +## waiting for remote KV and scheduled neither as new nor as cached, which routes it into the +## store scheduler's "pending load specs not yet scheduled" branch. That branch passes +## skip_save=None, so the rejected LoadSpec turns into a SAVE job, and _apply_current_save_block_ids +## finds no current block table for a request the core did not schedule: +## +## AssertionError: Missing current block table for store request ... +## +## EngineCore exits on the first such request. The assertion arrived with upstream #51358, first +## released in 0.29.0; 0.27.1 has the same branch without the assertion. +## +## Upstream fixed it in #54643 (34b1e9f7a6), released in 0.30.0: the branch skips a LoadSpec whose +## load this connector does not own. The hunk below is that change verbatim, rebased onto the 0.29.0 +## line numbers. It is also carried into new builds as patches/vllm/003_*.patch. + +RUN < Date: Thu, 24 Sep 2026 08:54:42 +0800 Subject: [PATCH 2/2] fix(vllm): carry the Mooncake store pending-load fix into new builds The post operation beside this commit repairs the published 0.29.0 images. A rebuild of 0.29.0 from the pack path would ship the defect again and overwrite that repair, so the same hunk goes where new builds are assembled. Both Dockerfile.vllm paths apply every patches/vllm/*.patch in order from the vLLM site-packages directory and stop on the first that fails, so the fix is one more file on each side, byte-identical between them as 001 and 002 already are. The hunk is upstream #54643 verbatim, with line numbers rebased onto v0.29.0. Verified with GNU patch 2.7.6 over a v0.29.0 tree in the order the build uses: 001, 002 and 003 all apply and the patched scheduler.py compiles. git apply --check accepts it against v0.29.0 and refuses it against v0.30.0, where the reverse check passes because upstream already carries the change. That refusal is deliberate and loud: moving the pack path to 0.30.0 or later has to delete this file, and a build that still carries it stops at the patch step rather than shipping silently. Signed-off-by: thxCode --- ...3_mooncake_store_rejected_pending_load.patch | 17 +++++++++++++++++ ...3_mooncake_store_rejected_pending_load.patch | 17 +++++++++++++++++ 2 files changed, 34 insertions(+) create mode 100644 pack/cuda/patches/vllm/003_mooncake_store_rejected_pending_load.patch create mode 100644 pack/rocm/patches/vllm/003_mooncake_store_rejected_pending_load.patch diff --git a/pack/cuda/patches/vllm/003_mooncake_store_rejected_pending_load.patch b/pack/cuda/patches/vllm/003_mooncake_store_rejected_pending_load.patch new file mode 100644 index 0000000..2d24c19 --- /dev/null +++ b/pack/cuda/patches/vllm/003_mooncake_store_rejected_pending_load.patch @@ -0,0 +1,17 @@ +diff --git a/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py b/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py +--- a/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py ++++ b/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py +@@ -371,7 +371,12 @@ + ) in self._unfinished_requests.items(): + if request_id not in request_ids and request_id not in cached_reqs.req_ids: + load_spec = self.load_specs.pop(request_id, None) +- if not load_spec: ++ # A load spec may have been proposed by this connector's ++ # lookup but rejected by MultiConnector in favor of another ++ # connector. Only the chosen connector may issue the pending ++ # load; the normal store path gets its blocks later from ++ # SchedulerOutput once the request is actually scheduled. ++ if load_spec is None or not load_spec.can_load: + continue + num_tokens_to_compute = load_spec.kvpool_cached_tokens + request_tracker = RequestTracker( diff --git a/pack/rocm/patches/vllm/003_mooncake_store_rejected_pending_load.patch b/pack/rocm/patches/vllm/003_mooncake_store_rejected_pending_load.patch new file mode 100644 index 0000000..2d24c19 --- /dev/null +++ b/pack/rocm/patches/vllm/003_mooncake_store_rejected_pending_load.patch @@ -0,0 +1,17 @@ +diff --git a/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py b/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py +--- a/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py ++++ b/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/store/scheduler.py +@@ -371,7 +371,12 @@ + ) in self._unfinished_requests.items(): + if request_id not in request_ids and request_id not in cached_reqs.req_ids: + load_spec = self.load_specs.pop(request_id, None) +- if not load_spec: ++ # A load spec may have been proposed by this connector's ++ # lookup but rejected by MultiConnector in favor of another ++ # connector. Only the chosen connector may issue the pending ++ # load; the normal store path gets its blocks later from ++ # SchedulerOutput once the request is actually scheduled. ++ if load_spec is None or not load_spec.can_load: + continue + num_tokens_to_compute = load_spec.kvpool_cached_tokens + request_tracker = RequestTracker(