CrossPool is a serving system for co-locating multiple SGLang models when user-driven KV Cache demand and model-driven FFN weight requirements do not line up.
KV Cache and FFN weights/execution are governed by different sizing axes:
| Resource | Main sizing driver | CrossPool treatment |
|---|---|---|
| Attention and KV Cache | Active requests, context lengths, and generation histories | SGLang keeps logical cache ownership; physical backing can be lent and reclaimed among Instances on one attention device. |
| FFN weights and execution | Model layer geometry, weight size, and FFN parallelism | FfnAgents retain model-specific true-TP shards on a shared FFN execution tier. |
A conventional co-located deployment binds these two axes to the same process and device reservation. One model may need large KV capacity while another mainly contributes resident FFN weights, yet each Instance must reserve both sides independently.
CrossPool separates the placement and sizing decisions without changing SGLang's request semantics. It shares FFN execution infrastructure and reallocates physical KV backing across co-located Instances. Logical KV contents and prefix caches remain isolated; CrossPool shares execution infrastructure and physical capacity, not cache contents.
The daemon control plane registers participants, coordinates SLO-aware KV capacity, and watches participant liveness. Each SGLang Instance Rank keeps attention and logical KV/prefix-cache ownership, then sends FFN work through a rank-local mailbox to its AtnAgent's Transport Kernel.
FfnAgents retain model-specific FFN weight shards as reusable GraphTemplates and replay them through per-lane CUDA Graphs. The Fabric carries the rank-local work between the AtnAgents and FfnAgents; physical KV backing can move between co-located Instances without sharing logical KV contents. See the system overview for complete process roles, request flow, and readiness contracts.
- Seamless SGLang integration: a pinned SGLang plugin installs architecture-specific adapters without replacing SGLang's request or KV Cache runtime.
- Independent resource placement: attention/KV and FFN/weight sides can be configured and sized independently for co-located models.
- SLO-aware elastic KV Cache backing: stable attention-side virtual addresses allow physical KV capacity to move between co-located Instances while SGLang keeps logical allocation and prefix-cache ownership.
- Layer-wise FFN execution: FfnAgents retain model-specific weight shards and execute model-defined FFN layers through GraphTemplates based on layer signatures.
- Linux on x86-64
- uv 0.12.17 or newer
- An uv-managed Python 3.12 interpreter
- CUDA Toolkit 13.2 and CCCL 3.2
- NVIDIA devices able to execute the selected kernels and CUDA Graphs, with CUDA IPC and NVSHMEM access required by the selected topology; the example uses one attention-side device and one FFN-side device
nvidia-cuda-mps-controlon PATH for daemon-owned attention MPS- Local model weights for serving and model-dependent validation
The native extension is built through uv and scikit-build-core, which obtains
suitable CMake and Ninja versions when needed. The system CUDA Toolkit provides
the native compiler and CCCL. CUDA bindings, Torch, SGLang, and the NVIDIA
NVSHMEM runtime are direct project dependencies. uv uses the interpreter pinned
in .python-version with managed Python downloads enabled. NVSHMEM runs through
the native C++/CUDA implementation; Python NVSHMEM bindings are not required.
Follow the two-device Qwen3-0.6B quick start to
configure a local checkpoint, start the four serving roles, send an HTTP
request through real FFN execution, and shut everything down in order.
The daemon starts and stops attention-side MPS; FFN processes run directly.
MPS endpoints use a private per-user root at /tmp/xpool-mps-<uid>/.
For the complete user-facing reference, see
Configuration. Start from
configs/xpool.example.toml and
.env.example; the Quick Start
shows a complete two-device setup.
CrossPool integrates with SGLang through architecture-discovered adapters. The
model architecture in config.json selects the adapter, while the configured
model ID resolves its weights. SGLang continues to own request scheduling,
attention, KV Cache, and output processing; CrossPool adds the shared FFN
execution path and elastic physical KV Cache backing. See
Supported Models for currently qualified model IDs.
Select xpool through SGLANG_PLUGINS and launch with
uv run xpool exec -- sglang serve .... The generic wrapper normalizes complete
deployment visibility and prepares attention MPS before the target imports its
runtime, preserving arguments, process identity and the inherited process group.
uv sync --group dev installs the private xpool-dev workspace member editably
alongside xpool, using the shared root lockfile and environment. Its console
entries are xtest.cli:main and xbench.cli:main; the production wheel contains
only production packages and its command.
xtest run is the canonical composition root. It runs native CTest,
Unit, Integration, and E2E stages in their accepted order, schedules device work
against explicit resource requirements, and retains artifacts under test-runs/
in the configured cache root (default .xpool-cache).
if [ -f .env ]; then export UV_ENV_FILE="$PWD/.env"; fi
# Complete resource-eligible suite.
uv run xtest run
# Selected canonical stages.
uv run xtest run --suite cext --suite integration
uv run xtest run --suite e2e --strict-requirementsSee tests/README.md for suite placement, requirements, and commands, and Test and Benchmark Tooling for process, device allocation, endpoint, and artifact ownership.
CMake uses ccache for C, C++, and CUDA when available and no compiler launcher
is already configured. To disable it for a build, add
--config-settings-package xpool:cmake.define.XPOOL_ENABLE_CCACHE=OFF
to the Quick Start sync command.
xbench measures multi-model serving through native SGLang streaming, either
against external endpoints or by starting a declared CrossPool deployment.
Each case fixes models, topology and workload; run retains measurements and
report generates offline latency distributions and throughput plots. See the
benchmark design for measurement and evidence
contracts, and the tooling tutorial for discovery,
execution, continuation and reports.
docs/designs/README.mdmaps the current implemented architecture.docs/plans/README.mdmaps candidate workstreams and their technical relationships.CONTEXT.mddefines the CrossPool domain language.src/xpool/contains configuration, runtime roles, the daemon, SGLang integration, and Python/native boundaries.src/cext-include/xpool/andsrc/cext/contain the C++/CUDA Transport and Fabric data plane.src/xpool-dev/contains the private development project and installedxkit,xtestandxbenchpackages for shared mechanisms, test tooling and benchmark tooling. Production and native sources retain their separate owners.tests/contains the native, Unit, Integration and E2E validation suites and their source-owned catalogue.benches/owns the benchmark catalogue, suite programs and inputs;configs/deployments/contains portable runtime scenes shared by the test and benchmark catalogues.docs/code-style.mddefines repository-wide coding conventions.
CrossPool is available under the MIT License.