Section 23

Recap

And further reading

A language model turns text into numerical representations, transforms them through its layers, and produces a probability for the next token. An inference engine repeats that computation while managing memory and sharing work across requests. The main ideas fit together as follows.

From text to a served response

  1. Text becomes a list of integer token IDstoken IDAn integer index into the vocabulary that uniquely identifies a token.See in glossary → via a tokenizertokenizerThe program that converts raw text into a sequence of integer token IDs (and back). Its vocabulary and merge rules are fixed before pre-training begins.See in glossary →, often a BPE-style tokenizer. Each ID indexes into that model’s vocabulary, whose size is model-specific.
  2. Each ID is mapped to a vector (“embeddingembeddingA dense vector representation of a token (typically d=2k–8k floats). Similar tokens get nearby vectors.See in glossary →”) of size dmodeld_{\text{model}} (4096-ish). That row is looked up from a giant matrix.
  3. Many current decoder-only LLMs inject position information with RoPERoPERotary Position Embeddings — rotates Q/K vectors by an angle proportional to position. Standard in modern LLMs.See in glossary → (rotating query/key vectors inside attentionself-attentionAttention where the queries, keys, and values all come from the same sequence, so each position can gather information from other allowed positions in that sequence.See in glossary →).
  4. A standard dense transformer layer has attention (cross-token mixing) and an MLPMLPMulti-Layer Perceptron — a stack of dense (matrix-multiply + nonlinearity) layers applied per-token. The transformer’s feed-forward block.See in glossary → (per-token nonlinear processing), supported by residual connectionsresidual connectionoutput = x + f(x). Lets gradients flow through deep stacks and means each block adds a refinement rather than rewriting.See in glossary → and normalization such as RMSNormRMSNormRoot Mean Square Normalization — a normalization layer that divides each activation by the root-mean-square (√(mean(x²))) of the whole vector, then multiplies by a learned per-dimension scale. Cheaper than LayerNorm (no mean subtraction, no learned bias) and empirically just as good. Standard in Llama-class models.See in glossary →.
  5. Attention computes scores QK⊤Q K^\top, applies softmax + causal mask, takes a weighted sum of VV. Multi-head splits this across many parallel “heads.”
  6. Stacking 32-128 of these blocks, plus an embedding lookup and an LM head, is the model.
  7. To generate, you take the final position’s logits, apply a sampling strategy (greedy / temperature / top-p), get a token, append, repeat.
  8. The first pass on the whole prompt (prefillprefillThe first forward pass that processes the entire prompt at once. Compute-bound, parallel over prompt tokens.See in glossary →) is often compute-bound for long or well-batched prompts. Small-batch single-token decodedecodeThe autoregressive phase: one forward pass per generated token. Memory-bandwidth-bound — the GPU mostly waits on weights.See in glossary → is commonly memory-bandwidth-bound because it repeatedly reads weights from HBMHBMHigh-Bandwidth Memory — the DRAM stack soldered next to the GPU die. H100 SXM has 80 GB at ~3.35 TB/s.See in glossary →.
  9. To avoid redoing prior token projections during decode, we cache the keys and values: the KV cacheKV cacheThe stored keys and values from all past tokens, so attention at step t only needs to compute Q for the new token.See in glossary →. It can be enormous; many modern models use GQAGQAGrouped-Query Attention — multiple query heads share one K/V head, shrinking the KV cache by 4–8× with minimal quality loss.See in glossary → to shrink it.
  10. Serving is constrained by HBM bandwidth and HBM capacity. The rest of the memory hierarchy (SRAMSRAMStatic Random-Access Memory — the on-chip scratchpad / L1+shared memory inside each SM. Tiny (~100s of KB per SM) but ~10× faster than HBM.See in glossary →, PCIe, NVLink, RDMA NIC) determines what kinds of parallelism work.
  11. Multiple users can share one GPU step via continuous batching: requests are admitted and completed dynamically as the scheduler has capacity.
  12. KV memory can be managed with PagedAttentionpaged attentionAn attention implementation that stores cached keys and values in fixed-size blocks, using a table to locate each sequence's blocks.See in glossary →: HBM is split into fixed-size blocks and each request has a block table, bounding per-request internal fragmentation.
  13. Pages can be shared across requests via prefix cachingprefix cachingSharing KV pages across requests that start with the same tokens (system prompts, few-shot prefixes), so the prefill is computed once.See in glossary →: identical completed prefix blocks can re-use the same physical KV blocks, reducing duplicated prefill and cache memory.
  14. Long prompts are processed via chunked prefill so they don’t block decoders.
  15. Speculative decodingspeculative decodingA small draft model proposes K tokens; the big target model verifies them all in one pass. Net effect: more tokens per target-model step.See in glossary → lets a draft model propose K tokens that the target verifies in one pass. Correct rejection sampling preserves the target distribution; the speedup depends on acceptance, draft cost, and serving conditions.
  16. Frontier models increasingly ship the draft model inside the model as multi-token-prediction (MTP) layers, then tune the whole stack for serving: GLM-5.2 shares the sparse-attention index and KV cache across draft steps, while Kimi K3 adapts its MTP layer into an EAGLE-style draft. Both optimize acceptance directly.
  17. Recurrent attention replaces a token-growing KV cache with fixed-size mutable state. Hybrid models must cache both that state and ordinary attention’s KV pages at exactly matching prefix boundaries.
  18. Million-token traffic makes cache locality and admission controladmission controlA serving policy that decides which requests may enter execution given available capacity. Budgeting request classes separately prevents very costly traffic from starving cheaper requests.See in glossary → fleet-level concerns: route sessions back to their cached prefixes and reserve separate capacity budgets so long requests cannot starve short ones.

When a single GPU isn’t enough, tensor parallelism (every matrix split across GPUs, all-reduce per layer over NVLink) and pipeline parallelism (layers split across GPUs, activations forwarded once per stage) carry the load, with data parallelism stacking replicas on top.

Underneath all of this, the same arithmetic ratio governs everything: how many FLOPs you do per byte of memory you read. Many serving optimizations improve that ratio by reusing data or doing more work each time it is read.

Further reading

These papers and implementations provide more detail on the algorithms and serving systems:

The papers

  • Vaswani et al., Attention Is All You Need (2017) — arXiv 1706.03762. The Transformer.
  • Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (2023) — arXiv 2309.06180. The vLLM paper.
  • Dao et al., FlashAttention (2022) — arXiv 2205.14135, and FlashAttention-2 (2023) — arXiv 2307.08691. Widely used attention kernels.
  • Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models (2022) — OSDI paper. Iteration-level scheduling.
  • Leviathan et al., Fast Inference from Transformers via Speculative Decoding (2023) — arXiv 2211.17192. An exact speculative-decoding protocol.
  • Chen et al., Accelerating Large Language Model Decoding with Speculative Sampling (2023) — arXiv 2302.01318. An independent formulation.
  • Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021) — arXiv 2104.09864. RoPE.
  • Ainslie et al., GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (2023) — arXiv 2305.13245. The K/V sharing method.
  • Cai et al., Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (2024) — arXiv 2401.10774.
  • Li et al., EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (2024) — arXiv 2401.15077.
  • Bai et al., IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse (2026) — arXiv 2603.12201. The IndexShare idea GLM-5.2 applies to its backbone and MTP layer.
  • Li et al., Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling (2026) — arXiv 2606.12370. Rejection sampling and the end-to-end TV loss for MTP drafts.
  • Z.ai, GLM-5.2: Built for Long-Horizon Tasks (2026) — blog post. The speculative-decoding stack from §21.
  • Kimi Team, Kimi K3: Open Frontier Intelligence (2026) — technical report. Recurrent-state caching, replay-based speculative decoding, and fleet scheduling from §22.

Codebases

Posts to read next

The vLLM scheduler, worker code, and attention kernels show how these ideas interact in a working server, including the bookkeeping and hardware-specific details that simplified examples omit.