CrossPool is a serving system for co-locating multiple SGLang models when user-driven KV Cache demand and model-driven FFN weight requirements do not line up.
KV Cache and FFN weights/execution are governed by different sizing axes:
| Resource | Main sizing driver | CrossPool treatment |
|---|---|---|
| Attention and KV Cache | Active requests, context lengths, and generation histories | SGLang keeps logical cache ownership; physical backing can be lent and reclaimed among Instances on one attention GPU. |
| FFN weights and execution | Model layer geometry, weight size, and FFN parallelism | FfnAgents retain model-specific true-TP shards on a shared FFN execution tier. |
A conventional co-located deployment binds these two axes to the same process and GPU reservation. One model may need large KV capacity while another mainly contributes resident FFN weights, yet each Instance must reserve both sides independently.
CrossPool separates the placement and sizing decisions without changing SGLang's request semantics. It shares FFN execution infrastructure and reallocates physical KV backing across co-located Instances. Logical KV contents and prefix caches remain isolated; CrossPool shares execution infrastructure and physical capacity, not cache contents.
The daemon control plane registers participants, coordinates SLO-aware KV capacity, and watches participant liveness. Each SGLang Instance Rank keeps attention and logical KV/prefix-cache ownership, then sends FFN work through a rank-local mailbox to its AtnAgent's Transport Kernel.
FfnAgents retain model-specific FFN weight shards as reusable GraphTemplates and replay them through per-lane CUDA Graphs. The Fabric carries the rank-local work between the AtnAgents and FfnAgents; physical KV backing can move between co-located Instances without sharing logical KV contents. See the system overview for complete process roles, request flow, and readiness contracts.
- Seamless SGLang integration: a pinned SGLang plugin installs architecture-specific adapters without replacing SGLang's request or KV Cache runtime.
- Independent resource placement: attention/KV and FFN/weight sides can be configured and sized independently for co-located models.
- SLO-aware elastic KV Cache backing: stable attention-side virtual addresses allow physical KV capacity to move between co-located Instances while SGLang keeps logical allocation and prefix-cache ownership.
- Layer-wise FFN execution: FfnAgents retain model-specific weight shards and execute model-defined FFN layers through GraphTemplates based on layer signatures.
- Linux on x86-64
- uv 0.12.17 or newer
- An uv-managed Python 3.12 interpreter
- CUDA Toolkit 13.2 and CCCL 3.2
- NVIDIA GPUs able to execute the selected kernels and CUDA graphs, with CUDA IPC and NVSHMEM access required by the selected topology; the example uses one attention-side GPU and one FFN-side GPU
- An externally managed CUDA MPS controller for runtime and GPU validation
- Local model weights for serving and model-dependent validation
The native extension is built through uv and scikit-build-core, which obtains
suitable CMake and Ninja versions when needed. The system CUDA Toolkit provides
the native compiler and CCCL. CUDA bindings, Torch, SGLang, and the NVIDIA
NVSHMEM runtime are direct project dependencies. uv uses the interpreter pinned
in .python-version with managed Python downloads enabled. NVSHMEM runs through
the native C++/CUDA implementation; Python NVSHMEM bindings are not required.
Follow the two-GPU Qwen3-0.6B quick start to configure a local checkpoint, start MPS and the four serving roles, send an HTTP request through real FFN execution, and shut everything down in order.
For the complete user-facing reference, see
Configuration. Start from
configs/xpool.example.toml and
.env.example; the Quick Start
shows a complete two-GPU setup.
CrossPool integrates with SGLang through architecture-discovered adapters. The
model architecture in config.json selects the adapter, while the configured
model ID resolves its weights. SGLang continues to own request scheduling,
attention, KV Cache, and output processing; CrossPool adds the shared FFN
execution path and elastic physical KV Cache backing. See
Supported Models for currently qualified model IDs.
xtest run is the canonical composition root. It runs native CTest,
Unit, Integration, and E2E stages in their accepted order, schedules GPU work
against explicit resource requirements, and retains artifacts under
.xpool-cache/test-runs/.
if [ -f .env ]; then export UV_ENV_FILE="$PWD/.env"; fi
# Complete resource-eligible suite.
uv run xtest run
# Selected canonical stages.
uv run xtest run --suite cext --suite integration
uv run xtest run --suite e2e --strict-requirementsSee tests/README.md for suite placement, requirements, and commands, and Test and Benchmark Tooling for process, GPU lease, endpoint, and artifact ownership.
CMake uses ccache for C, C++, and CUDA when available and no compiler launcher
is already configured. To disable it for a build, add
--config-settings-package xpool:cmake.define.XPOOL_ENABLE_CCACHE=OFF
to the Quick Start sync command.
xbench measures multi-model LLM serving through native SGLang streaming. It
supports client-only load against externally owned endpoints and owned execution
that starts the declared CrossPool model combination and topology. Prompt JSONL
or random token prompts combine independently with trace JSONL or Poisson arrivals.
Random prompts require matching local tokenizer/model metadata; inputs are prepared
offline and retained for replay.
run computes request metrics and retains them with request/event JSONL and
execution checkpoints. report aggregates saved metrics and event samples into
distributions and logical input/output throughput, then exports CSV and
paper-layout PDF/SVG/PNG figures inside each repetition's report/ directory.
Run inputs expand to independent repetition reports. Repeated reporting
overwrites generated files and preserves unrelated files and the measurement.
if [ -f .env ]; then export UV_ENV_FILE="$PWD/.env"; fi
uv run xbench list
uv run xbench run --case serving-001
uv run xbench report .xpool-cache/bench-runs/RUN_ID
uv run xbench clean --dry-runThe checked-in catalogue uses the
two-Qwen deployment,
one attention GPU, one FFN GPU, external MPS and local checkpoints resolved from
XPOOL_CONFIG. Owned cases inherit machine paths and runtime policy from that
complete configuration while their portable deployment supplies topology and
SLO; optional runtime_config selects an explicit base.
Supply --catalog FILE for other scenarios; relative input paths resolve against
that catalogue. Each catalogue names a source module below its sibling suites/
directory; that program orchestrates the scenario through installed tooling.
Client catalogues declare external endpoints instead of an owned deployment. An
installed invocation outside this checkout supplies an explicit catalogue.
The root uv sync --group dev installs the private xpool-dev workspace member
editably alongside production xpool, using one root lockfile and virtual
environment. Its direct console entries are xtest.cli:main and
xbench.cli:main; the released xpool wheel contains production code and the
xpool command. Both development tools expose list/run/report/clean; xtest
inventory and execution require source suites, while offline reports work outside
a checkout.
See Test and Benchmark Tooling for ownership, dataset
and metric contracts. Performance values are report-only, not readiness gates.
docs/designs/README.mdmaps the current implemented architecture.docs/plans/README.mdmaps candidate workstreams and their technical relationships.CONTEXT.mddefines the CrossPool domain language.src/xpool/contains configuration, runtime roles, the daemon, SGLang integration, and Python/native boundaries.src/cext-include/xpool/andsrc/cext/contain the C++/CUDA Transport and Fabric data plane.src/xpool-dev/contains the private development project and installedxkit,xtestandxbenchpackages for shared mechanisms, test tooling and benchmark tooling. Production and native sources retain their separate owners.tests/contains the native, Unit, Integration and E2E validation suites and their source-owned catalogue.benches/owns the benchmark catalogue, suite programs and inputs;configs/deployments/contains portable runtime scenes shared by the test and benchmark catalogues.docs/code-style.mddefines repository-wide coding conventions.
CrossPool is available under the MIT License.