The cost of moving KV
Bytes, bandwidth, and topology
The prefill worker has completed an 8,192-token prompt. How much data must move before a fresh decode worker can use its state? The answer starts with model geometry, not with the number of visible words in the question.
For a conventional full-attention model using grouped-query attention, assume the cache stores one key and one value per KV head at every layer and position. Ignoring allocator padding and metadata, the payload is:
Here, is the layer count, the number of KV heads, the head dimension, the number of cached positions, and the bytes per stored element. The leading two counts K and V. This is the same cache geometry introduced in the KV cache chapter, now used as a network budget.
Work the numbers all the way through
Use a model geometry with 32 layers, 8 KV heads, head dimension 128, and two-byte elements. One token requires:
An 8,192-token prompt therefore occupies 1,073,741,824 bytes: exactly 1 GiB in this simplified representation. A 32,768-token prompt occupies 4 GiB. The tiny question appended to a long handbook can require a large handoff because the decoder needs the whole relevant history, not just the newly typed suffix.
Be consistent about units. A gigabyte, GB, is one billion bytes. A gibibyte, GiB, is bytes. Network specifications frequently use bits per second; divide by eight to convert a bit rate to a byte rate before doing the calculation. A nominal 200 Gb/s link has a 25 GB/s raw byte-rate ceiling, before overhead and bottlenecks elsewhere in the path.
With effective bandwidth , the payload-only transfer lower bound is:
At 25 GB/s, transferring 1 GiB takes at least about 42.95 milliseconds. At 5 GB/s, it takes at least 214.75 milliseconds. Neither number includes allocation, registration, synchronization, staging, queueing, or conversion. These are arithmetic examples, not measured SGLang transport latencies.
A fast network label is not an end-to-end rate
The path may cross GPU memory, a PCIe connection, a network interface, a fabric, another interface, and a destination memory path. Traffic from tensor parallelism, other transfers, or storage can share parts of that route. Effective bandwidth is constrained by the whole path and by concurrent users of it.
Remote direct memory accessRDMARemote DMA — letting one node’s NIC write directly into another node’s memory without involving the CPU. The basis of InfiniBand and RoCE.See in glossary →, or RDMA, supports movement between registered memory regions with less host involvement in the data path. It does not provide infinite bandwidth or eliminate registration, topology, and completion management. Supported direct GPU paths can avoid costly staging copies, but their availability depends on hardware and transport configuration.
The Mooncake paper treats communication and cache placement as central serving concerns. Its broader architecture is useful context for why moving reusable state can be valuable and why scheduling those movements matters. The transfer backend used by SGLang is one component of that larger design space. Mooncake paper, v4
A cache hit can still require the full transfer
Suppose 8,000 of the 8,192 prompt tokens are already cached on the prefill worker. The worker computes only the missing suffix, subject to execution-boundary details. If the chosen decode worker has none of that state, it still needs the required prompt KV.
The compute savings scale with the missing portion. Transfer savings depend on the destination’s usable state and the transfer protocol. Do not subtract the prefill hit length from the network payload merely because the source avoided recomputing it.
If a supported destination-cache path already has a compatible prefix, the protocol may be able to avoid transferring some bytes. That requires actual implementation support and an established matching state. Our workbench assumes a fresh destination and transfers the full conventional prompt cache once. It intentionally keeps prefill reuse separate from destination reuse.
Parallelism changes layout as well as capacity
With tensor parallelism, cache values may be partitioned or replicated across ranks according to model geometry and execution strategy. “Divide by the number of GPUs” is not a universally valid payload rule. A sender rank must put the correct values into the layout that its receiving ranks expect.
Using different tensor-parallel sizes for prefill and decode introduces an extra mapping problem. Data that is contiguous in one layout can be divided differently in the other. SGLang documents a Mooncake staging-buffer path for supported heterogeneous-TP, non-MLA configurations. Its gathering and scattering work addresses that mismatch; it also consumes buffers and bandwidth. Heterogeneous TP configuration
Our formula assumes conventional dense KV. Multi-head latent attention, sliding-window layers, recurrent state, sparse attention, and draft-model state can produce different payloads. Count the representation actually handed off. A model with fewer KV bytes may still spend time converting layouts or moving additional state not included in a simple K-plus-V estimate.
Quantization is not free bandwidth
If every KV element changes from two bytes to one, the idealized dense payload halves. That is a useful first estimate. Actual formats can require scale metadata, alignment, conversion, and compatible kernels. The model’s output quality and numerical behavior also need evaluation.
The workbench’s one-byte option changes payload size only. It does not assert that the selected model supports that configuration or that its compute rate remains unchanged. SGLang’s quantized-KV documentation is the place to check supported modes for a real deployment. Quantized KV cache
Overlap changes the critical path
If a transfer starts only after all prompt processing ends, prefill and transfer times add on that request’s critical path. A supported pipelined transfer can begin moving completed portions while later computation continues. Perfect overlap would hide some transfer time, but the destination still cannot consume missing required state.
Layerwise or chunkwise transfer is therefore a scheduling optimization as well as a bandwidth optimization. Smaller pieces can start sooner but incur more per-transfer bookkeeping. Contention can also reduce the bandwidth seen by the overlapping operations. Verify the actual backend path instead of assuming all deployments overlap all layers automatically.
Budget the whole service
The workbench combines the payload calculation with a deliberately simple capacity model. Eight workers are split between prefill and decode. Each prefill worker processes an assumed 8,000 input tokens per second; each decode worker processes an assumed aggregate 800 output tokens per second. A single shared network bottleneck carries all handoffs.
Budget a disaggregated service
Dense GQA example: 32 layers, 8 KV heads, head dimension 128. Transfer the full prompt once. Fixed illustrative aggregate rates: 8,000 input tokens/s per prefill worker and 800 output tokens/s per decode worker. Quantization changes payload only here; scales, conversions, topology, and rate changes are omitted.
- KV payload
- 1.000 GiB
- Payload transfer lower bound
- 42.95 ms
- Pool split
- 3 prefill / 5 decode
- Shared link offered load
- 8.6%
| Stage | Capacity bound (requests/s) |
|---|---|
| Prefill | 2.93 |
| Decode | 7.81 |
| Shared network | 23.28 |
Bottleneck bound: 2.93 requests/s. Offered load is below this bound; latency objectives still need measurement.
- 1. Prefill
1024 ms - 2. Transfer
≥ 42.95 ms - 3. Decode
continues after readiness
The illustrative timeline excludes queueing and assumes serialized prefill and transfer. Aggregate output-token capacity does not determine an individual request's inter-token latency.
Illustrative model, not a SGLang benchmark. Controls update immediately; no automatic animation.
For fixed input and output lengths, dividing each pool’s token rate by the relevant length gives a requests-per-second capacity bound. The network bound is bytes per second divided by bytes per request. The smallest bound identifies the first stage that cannot sustain additional arrivals under these assumptions.
At the default 8,192 input tokens, 512 output tokens, and three prefill workers, prefill capacity is about 2.93 requests per second. Five decode workers provide about 7.81 requests per second. A 25 GB/s shared link provides about 23.28 transfers per second. This toy deployment is prefill-limited; buying more network bandwidth would not remove its first bottleneck.
Now increase output length or shift more workers into prefill. The limiting stage changes. Raising the arrival rate above the minimum bound predicts unbounded queue growth in an idealized indefinitely running system without rejection. Staying below it does not guarantee a latency objective: burstiness, memory limits, variable service times, and batching remain outside this calculation.
This budget is a screening tool. It can rule out impossible expectations and identify which measurement is worth taking next. It cannot predict a real deployment’s exact goodput from a handful of sliders.
Sources and further reading
- Mooncake v4: KV-centric communication and placement.
- SGLang PD documentation: transfer backends, topology-related options, and heterogeneous TP.
- Quantized KV cache: supported representation choices.
- NIXL repository and Mooncake repository: primary transport references. Online sources checked September 24, 2026.