Section 09

Routing and balancing the service

Cache affinity meets queue pressure

Two prefill workers can answer the next handbook question. Worker A already has the document’s KV, but a queue of questions is waiting there. Worker B is idle, but its cache is cold. Which destination gives the user the better result?

Routing by cache affinity alone says A. Routing by immediate load alone says B. A useful decision considers both the work saved by a hit and the delay incurred while waiting for that hit. Once prefill and decode are separate, there is a second destination to choose and a transfer path connecting them.

Cache affinity has a waiting-time price

Suppose worker A would need only 20 milliseconds of additional prefill after reaching the front of its queue. Worker B would need 200 milliseconds from scratch. If A’s queue adds 40 milliseconds, A looks attractive. If it adds 600, B may finish much sooner despite repeating work.

A rough teaching score is:

estimated finish=estimated queue wait+remaining service time.\text{estimated finish} = \text{estimated queue wait} + \text{remaining service time}.

This is not SGLang’s exact routing formula. It explains why cache match and load must be evaluated together. For PD, you would also care about destination capacity and the transfer path. The fastest prefill completion is not necessarily the fastest usable streamed answer.

SGLang’s gateway exposes cache-aware routing with load-balancing thresholds and supports separate policies for prefill and decode. Its documented PD example pairs cache-aware prefill routing with power-of-two decode routing. The available knobs express a tradeoff between locality and load distribution. Model Gateway documentation

The router’s tree is not the worker’s KV

A gateway may track which prompts have been routed to which workers and use that information to estimate likely reuse. The worker’s cache manager owns the actual cache entries, allocations, and eviction decisions. Those are different views of state.

A worker may have evicted a prefix after the router last saw a related request. It may have restarted. A storage fetch may be pending rather than complete. A routing decision based on likely locality should therefore not be described as a guarantee that every byte is resident and ready.

Multiple gateway replicas introduce another source of uncertainty. Their routing histories can differ unless a specific mechanism synchronizes them. The online gateway documentation explicitly discusses unsynchronized cache trees across router replicas. We should expect the consistency of the routing view to affect locality, rather than assuming a single globally exact radix tree. Gateway deployment guidance

For the assistant, this means cache-hit measurements belong at the workers as well as the gateway. If routing predicts a warm prefix but the worker recomputes it, examine eviction, restarts, namespaces, and request formatting before blaming the attention kernel.

Decode routing has a different objective

At a prefill worker, a long matching prefix can directly remove prompt computation. At a decode worker, the immediate concern is often whether incoming KV and future output will fit while existing requests keep acceptable token cadence.

Counting active requests gives only a rough estimate. Ten short-context requests about to finish can be easier than five long-context requests generating lengthy reports. Their memory footprints and remaining service times differ. Output length is uncertain unless the application imposes a fixed cap, and even a cap is not a prediction of actual generation length.

A power-of-two policy samples two candidates and prefers the less loaded one according to its load signal. It is a practical balancing technique, not clairvoyance about future output. Report what the chosen load metric actually measures: request counts, token counts, queue length, or another quantity.

Request placement can also affect locality on the decode side where supported. But the basic PD walkthrough assumes transferred state reaches the chosen destination; it does not presume that arbitrary decode workers already share a coherent cache. Any additional reuse path should be verified for the selected backend and model.

Size pools in units of work

Suppose arrivals average four requests per second, each with 8,000 uncached input tokens and 500 output tokens. The offered work is 32,000 input tokens per second and 2,000 output tokens per second. If a measured prefill worker sustains 10,000 relevant input tokens per second, four workers are a rough minimum before allowing for bursts and headroom. If a decode worker sustains 800 output tokens per second under the target latency, three are similarly a rough minimum.

This calculation uses measured rates at the desired operating point. A maximum-throughput decode rate from a huge batch may not be usable under a strict token-latency target. Likewise, a prefill rate measured with short prompts may not apply to long-context attention.

Now let prefix reuse remove half the input work. The prefill requirement might drop substantially, but the input histories can remain large and the output demand is unchanged. The network and decode-memory budgets do not necessarily fall by half. Separate accounting for input work, transferred bytes, and active state prevents an attractive cache-hit graph from masking a different bottleneck.

Pools need headroom, not just average balance

A system with mean arrivals equal to mean capacity has no spare service budget to drain ordinary fluctuations. A short burst can create a queue that takes a long time to clear. Long requests and correlated arrivals make this worse.

Consider a company-wide email announcing a new handbook. Many users open the assistant at once. The first requests may all be cold; subsequent requests can benefit from reuse. Average daily hit rate hides the demanding first minute. Load tests should include that transition, not just a steady warm state.

Scaling a pool changes more than its worker count. A new worker starts with cold local caches, requires model initialization, and may change routing affinity. Moving all traffic immediately to an empty worker can create a temporary prefill surge. A useful rollout observes warm-up and latency before declaring the capacity available.

Admission control prevents stranded work

When decode is full, allowing the prefill pool to run indefinitely creates state that waits for a destination. The preallocation and transfer queues from chapter 7 provide concrete places where pressure appears. Their occupancy is part of the service’s resource budget.

BackpressurebackpressureLimiting upstream admission or production when a downstream stage lacks capacity, preventing unbounded queues.See in glossary → means downstream capacity limits how quickly upstream stages admit or produce work. It can take the form of bounded queues, delayed dispatch, or rejection when the service cannot meet its policy. The goal is to prevent one saturated stage from converting every other stage into an ever-growing buffer.

Mooncake studies overload explicitly, including admission decisions aimed at preserving useful throughput under latency constraints. That is a research perspective on why admitting every request is not always equivalent to serving every user well. It does not imply that SGLang automatically implements Mooncake’s complete prediction and rejection policy when its transfer engine is enabled. Mooncake paper

If requests are rejected, count them. A latency report over only admitted requests can look excellent while most users receive errors. Report offered rate, admitted rate, completion rate, and goodput together. The same principle applies to automatic retries: they add internal load even when the external request count stays fixed.

Topology can make a good pair a bad route

A prefill worker and decode worker may each be lightly loaded but connected by a congested path. Another pair may have slightly more compute work and a much better handoff. DistServe’s placement analysis highlights this coupling between stage assignment and bandwidth. DistServe

Large mixture-of-experts deployments add another interaction. Attention and expert computation can use different parallel layouts, and collective communication may share infrastructure with KV movement. SGLang’s Kimi K2 deployment report combines PD with large-scale expert parallelism. Read its hardware, model, and traffic conditions alongside its results; its chosen configuration is a case study, not a recipe for an eight-billion-parameter document assistant. Kimi K2 deployment report

The service view

For our assistant, a useful dashboard separates prefill queueing, cache reuse, transfer occupancy, decode admission, and token cadence. Add worker restarts and gateway retry counts. These signals explain whether a slow answer needs a warmer route, more prefill capacity, more decode capacity, or less traffic admitted to an already full path.

The next chapter turns these ideas into a small deployment and a controlled experiment. The purpose is to produce evidence for an architectural choice, not to find a single flag that makes every workload fast.

Sources and further reading