What SGLang optimizes
From language model programs to a serving engine
A document assistant receives a 12,000-token company handbook followed by a 40-token question. It answers. A second employee sends the same handbook and a different question. A third asks for three independent analyses of one paragraph. Meanwhile, the first employee asks a follow-up.
An inference engine sees requests, token sequences, and batches. The application sees a document, several readers, and conversations that branch. Much of the arithmetic is repeated because those requests have related beginnings. How much can a runtime discover and reuse? Once reuse makes the GPU’s job smaller, can the CPU prepare work quickly enough? And when one machine is full, where should the next request go?
Those questions lead naturally into SGLangSGLangAn inference framework whose original design paired a language for structured model programs with a runtime that reuses and schedules their computation.See in glossary →. Its original design combined a language for expressing model calls with a runtime that exploited the relationships between them. The paper calls these applications language model programs: ordinary control flow interleaved with generation, branching, and structured inputs or outputs. That is the historical starting point for this explainer. SGLang paper, §§1–2
Where this companion begins
You should already recognize tokens, attention, model weights, and the KV cache from LLM Inference. The prefill/decode chapter and KV cache chapter are useful refreshers. We will use their vocabulary without rebuilding the transformer layer by layer. Experience deploying a model is not required.
Our subject is the serving system around that model: the decisions that determine whether useful work is repeated, delayed, transferred, or discarded. We will follow the handbook assistant from one instance to a service with multiple prefill and decode workers. Numbers in worked examples are chosen to expose a tradeoff; they are not performance claims about a particular accelerator.
One boundary matters throughout. Improving a service can mean making one answer arrive sooner, allowing more users to receive acceptable answers, or using fewer resources for the same demand. Those objectives overlap, but a change that improves one may harm another. A bigger batch can increase total output while making an individual request wait. A remote cache can avoid computation while adding transfer latency.
A frontend and a runtime
In the original programming model, a developer can append text, request generation, fork a prompt into several continuations, and join their results. These primitives make relationships between calls explicit. The paper’s runtime can also function independently of that frontend; using its inference optimizations does not require rewriting an application as a SGLang program. SGLang paper, §2
For our assistant, imagine the following application-level workflow. This is pseudocode, not a current frontend API example:
base = system_instructions + handbook
answers = parallel_generate(
base + "Explain the travel policy",
base + "Explain the expense policy",
base + "List unresolved questions"
)
summary = generate(base + answers + "Combine these findings")
The three branches share the same initial token sequence. The final call may share that beginning too, even though it has a much longer suffix. A runtime that retains earlier KV can exploit those shared beginnings without knowing what a handbook means.
Notice what cannot be reused automatically. If a branch paraphrases the document, moves it after the question, or uses a different chat template, the prefix is different. Similar meaning is insufficient. The useful relationship is between exact computational histories, including the model and cache-affecting settings.
Follow one request
A useful first map separates CPU coordination from GPU execution. The concrete process boundaries vary with configuration, but these responsibilities provide a stable way to read SGLang’s runtime code. Its scheduler owns batching and coordinates work with the model worker; tokenization and detokenization are surrounding serving responsibilities. SGLang v0.5.20 scheduler
- 1. API and tokenization
Validate the request and turn its formatted prompt into token IDs. - 2. Cache and scheduler
Match a reusable prefix, reserve memory, and choose a batch. - 3. Model execution
Process missing prompt tokens, then generate continuations. - 4. Output and cache lifetime
Stream text, release active references, and retain eligible KV.
Suppose 12,000 of 12,040 prompt tokens are already cached. The scheduler still has work to do: locate those entries, confirm they are usable, allocate space for the suffix and generation, and assemble model inputs. The GPU still processes the uncached suffix. It then generates the answer, repeatedly attending to previous state. A large hit removes repeated prompt work; it does not turn generation into a lookup of an old answer.
This distinction gives us two independent measures. Cached input tokens describe avoided prompt processing. Output tokens per second describe ongoing generation. They can move differently. A service may become much faster to first token while its subsequent token cadence barely changes.
Three places to save work
The first opportunity is reuse within an instance. A shared prefix is already in that instance’s GPU memory, and a new request can refer to it. Chapter 2 explains the radix tree used to organize those prefixes.
The second is reuse across memory tiers. An old prefix has left GPU memory but may still be in host memory or storage. Retrieving it spends bandwidth to avoid recomputation. Chapter 5 follows this choice through HiCache.
The third is specialization across workers. A prefill worker processes a prompt and transfers the resulting state to a decode worker. Both workers participate in the same live request. This handoff is required by that deployment even for a completely new prompt with no reusable prefix. Chapters 6–9 explain the benefits and costs.
Consider a cold request that shares nothing with any earlier user. Radix reuse offers no savings. A storage cache also misses. Yet separating its prefill from other users’ decoding might still improve their token cadence. Conversely, a small colocated server with high prefix reuse may gain little from splitting itself into separate worker pools. These mechanisms solve different problems and should be evaluated separately.
What changes relative to the vLLM story?
The vLLM companion introduced physical block allocation, prefix caching, continuous batching, and speculative decoding. Keep those concepts. This explainer changes the point of view: we begin with relationships between requests, then follow how scheduling and placement interact with reuse.
A radix tree and a page allocator operate at different levels. One identifies a reusable history; the other manages physical storage. An inference engine can use both. Similarly, using a router does not mean every worker shares one cache, and separating prefill from decode does not automatically pool their memory. We will keep asking which process owns the bytes and which process merely knows where they might be.
By the end, you should be able to draw the path of the handbook’s KV cache, name the queues a question can wait in, and estimate the bytes that cross a network when the request moves between workers. Those are practical tools for understanding why an inference system is fast—or why it slows down when the workload changes.
Sources and further reading
- SGLang: Efficient Execution of Structured Language Model Programs, v2: the original frontend/runtime design and evaluation.
- SGLang v0.5.20: the implementation snapshot used here.
- Scheduler source at v0.5.20: the central coordination code; read for responsibilities before reading individual branches.