Putting the system together
Deploy, measure, and choose
The handbook assistant has accumulated several possible optimizations: stable prefixes, a radix cache, overlap scheduling, a larger cache hierarchy, and separate prefill and decode pools. Each has a plausible reason to help. A deployment decision still needs a measured comparison with a simpler baseline.
This chapter provides a small PD walkthrough and an experiment design. The commands follow the release-pinned SGLang v0.5.20 documentation. They have been checked against those sources, not executed on GPU hardware for this explainer. The interactive models elsewhere in the companion are likewise educational calculations rather than benchmark results.
Establish one reproducible environment
Use a Linux NVIDIA host with two supported CUDA GPUs, enough memory on each for the selected model and its cache, a compatible driver, and Docker with GPU support. For this example, assume two 24 GiB-or-larger GPUs, modest concurrency, and access to the Llama 3.1 8B Instruct weights. Actual memory needs depend on configured lengths, buffers, and backend behavior.
Use the same model revision, tokenizer, dtype, and runtime environment on both workers. The versioned image bundles the runtime, gateway build, and transfer dependencies; the pinned Dockerfile documents those components. This release uses a CUDA 13 environment, so use its installation requirements rather than older CUDA 12 instructions. Release notes, pinned Dockerfile
Start a container with the two GPUs visible. HF_TOKEN below refers to an already configured environment variable; it is not a literal credential. The model cache volume avoids repeated weight downloads.
# Run on the Linux GPU host.
docker run --rm -it --name sglang-pd-demo \
--gpus all --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
--env HF_TOKEN \
lmsysorg/sglang:v0.5.20 bash
Keep this container open. Use additional host terminals with docker exec -it sglang-pd-demo bash for each process below. All processes then share the container’s network namespace and installed packages. For strict reproduction, record the resolved image digest and model commit as well as the human-readable tags.
Launch one worker for each phase
In the first container shell, start the prefill worker on GPU 0:
python3 -m sglang.launch_server \
--model-path meta-llama/Llama-3.1-8B-Instruct \
--disaggregation-mode prefill \
--disaggregation-transfer-backend nixl \
--base-gpu-id 0 \
--port 30000
In a second container shell, start decode on GPU 1:
python3 -m sglang.launch_server \
--model-path meta-llama/Llama-3.1-8B-Instruct \
--disaggregation-mode decode \
--disaggregation-transfer-backend nixl \
--base-gpu-id 1 \
--port 30001
Wait for both workers to finish initialization. In a third container shell, start the gateway:
python3 -m sglang_router.launch_router \
--pd-disaggregation \
--prefill http://127.0.0.1:30000 \
--decode http://127.0.0.1:30001 \
--host 0.0.0.0 \
--port 8000
These commands adapt the documented single-node NIXL example. They demonstrate role assignment and one public endpoint, not a production pool ratio. The transport still requires a working supported device path. If bootstrap succeeds but transfer stalls, inspect the backend and worker logs rather than assuming the HTTP gateway is the bottleneck. Pinned PD setup
A two-GPU demonstration is not evidence that two role-specific workers outperform two colocated replicas. That equal-resource comparison belongs in the benchmark.
Send a streamed request
From the host, send a request to the mapped gateway port:
curl -N http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [
{"role": "system", "content": "Answer briefly using the supplied handbook."},
{"role": "user", "content": "Handbook: Expenses require receipts. Question: What must I attach?"}
],
"temperature": 0,
"max_tokens": 64,
"stream": true
}'
The tiny prompt verifies the request path. It does not exercise the long-prefix workload that motivated the architecture. Check that the stream completes and worker logs show the intended roles. A response alone does not establish that every desired optimization is active.
For multi-turn tests, send explicit conversation history consistently. Process-local response history is not automatically shared between PD workers. The PD documentation describes limitations on stateful Responses API workflows; ordinary chat history supplied in the request avoids relying on that separate feature. PD API behavior
Begin with a simple load test
Run a benchmark inside another container shell after all processes are ready:
python3 -m sglang.bench_serving \
--backend sglang \
--host 127.0.0.1 \
--port 8000 \
--model meta-llama/Llama-3.1-8B-Instruct \
--dataset-name random \
--random-input-len 4096 \
--random-output-len 256 \
--num-prompts 200 \
--max-concurrency 8
This is a smoke benchmark for the pipeline. Random prompts do not represent the handbook’s stable shared prefixes. Inspect the generated length distribution and stopping behavior before interpreting the output; requested maximum output length is not always the number of tokens actually emitted. SGLang distinguishes online serving benchmarks from lower-level kernel and offline throughput tools. Benchmark and profiling guide
For the cache study, build a fixed request trace with explicit shared token prefixes, unique suffixes, and a documented warm-up phase. Preserve that trace across configurations. Report whether template tokens are included in input lengths and whether generated output is allowed to terminate early.
Run a controlled matrix
Start with a colocated server, then two colocated replicas using the same two GPUs, then the one-prefill/one-decode setup. Compare configurations at equal total resources. Tune the colocated chunking baseline before drawing conclusions about disaggregation.
| Variable | Cases to include | What it reveals |
|---|---|---|
| Prefix residency | Cold; warm GPU; host/storage where configured | Compute savings versus restoration cost |
| Input length | 1k, 8k, 32k tokens within supported limits | Prefill work and KV payload growth |
| Output length | Short answers; long reports | Decode capacity and state lifetime |
| Arrival pattern | Low load; steady sweep; burst | Queueing, headroom, and overload |
| Shared prefix | None; handbook shared; mixed documents | Locality and cache churn |
| Deployment | Colocated; chunked; PD | Isolation benefits versus handoff costs |
| Pool ratio | Several splits at a larger fixed GPU budget | Stage bottlenecks and stranded capacity |
Do not change all variables at once. First establish whether the workload is prefill-, decode-, memory-, or transfer-limited. Then test the mechanism intended to address that bottleneck. Repeat representative runs to distinguish a stable effect from warm-up noise or a temporary placement advantage.
Measure what the user sees
Capture client arrival, first-token, subsequent-token, and completion timestamps. Report TTFT percentiles, ITL distributions, per-request TPOT, completed output tokens per second, errors, and goodput under explicitly stated objectives. Keep offered and admitted request rates visible. If requests time out, they belong in the results.
For each worker, observe running and queued requests, cache occupancy and reuse, prefill/decode work, and transfer waits where exposed. SGLang’s production metrics and request-tracing documents describe available instrumentation. Metric names and semantics should be checked against the pinned runtime rather than copied from an unrelated dashboard. Production metrics, request tracing
Profile prefill and decode workers separately when investigating PD execution. A client-level latency spike can come from either worker or from the interval between them. Align traces by request identity and timestamps; a fast prefill kernel does not explain a long destination-allocation wait.
Treat initialization and cache warming deliberately. Kernel capture, compilation, weight loading, and storage warm-up can distort a short run. Report both cold-start behavior where relevant and steady-state behavior after a defined warm-up. Avoid quietly dropping inconvenient requests from the measurement window.
Make a workload-based choice
If one instance meets the required goodput, a simpler colocated service may be the right starting point. If repeated prompt processing dominates, verify prefix structure and cache effectiveness. If reusable state spills out of GPU memory, evaluate host or storage restoration. If prefills disrupt token cadence despite suitable chunking, test PD and its transfer budget.
Keep the architecture only if the measured benefit justifies its extra resource and operational costs. Record the configuration, workload, SLO definition, and rejected or failed requests alongside any headline improvement. That record makes the result useful when the handbook becomes longer, the model changes, or users begin requesting reports instead of short answers.
Annotated reading guide
- SGLang, v2: read §§2–4 for language model programs, radix reuse, and constrained decoding. Its experimental comparisons describe the paper’s historical evaluation.
- SGLang v0.4 scheduler discussion: explains CPU/GPU overlap and cache-aware balancing as later runtime developments.
- HiCache design: separates local residency, storage lookup, and transfer policy.
- DistServe, v3: develops the case for phase separation under latency objectives and bandwidth constraints.
- Mooncake, v4: treats KV placement, communication, and overload as parts of one serving problem. Its full system is broader than its transfer-engine integration.
- Kimi K2 deployment report: shows PD and expert parallelism together on a specific large cluster; preserve those conditions when discussing its results.
- Release-pinned source tree: the reference for code and commands in this companion. Online documentation was checked September 24, 2026 and may evolve after that date.
You can now follow the same state from a shared prefix, through a scheduler and cache hierarchy, across a transfer boundary, and into a decode batch. When a request is slow, that path tells you which question to ask next: what was recomputed, what was waiting, what was moved, and which resource could not keep up?