fix(vllm): support synchronous GPUDirect loads - #356
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
load_async=falsefor the vLLM connectorstart_load_kv, before the request forward passWAITING_FOR_REMOTE_KVSpath unchanged forload_async=trueProblem
The async path issues remote GPU writes from
get_finished()after model compute has already launched. On hybrid state-cache models, vLLM does not currently expose a per-block ownership fence proving those destination blocks are disjoint from concurrent recurrent-state compute.On GLM-5.3-Flash TP8, 64K-input C10 external-cache replay consistently corrupted in-flight CUDA state:
The native client reported exact-length GET hits and zero I/O errors, so retrying transport or changing node dedup did not address the ownership race.
Validation
Hardware: xb01-0064, 8x B200. Model: GLM-5.3-Flash TP8+EP, no MTP. Workload: 100 prompts, 65,536 input tokens, 500 output tokens, concurrency 10.
With
load_async=false:Tests: