Optimizing speculative decoding
Index sharing, acceptance-aware losses, and built-in draft models
Sources: GLM-5.2: Built for Long-Horizon Tasks — Z.ai, 2026; IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse — Bai et al., 2026; Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling — Li et al., 2026
Speculative decoding can use a separate small model, an EAGLE draft component, or Medusa heads. Some open models include draft layers trained alongside the main model. Multi-Token PredictionMulti-Token PredictionMulti-Token Prediction (MTP) — a training objective where the model predicts several future tokens at each position (not just the next one), densifying the learning signal and enabling faster speculative decoding later.See in glossary → layers, trained during pre-training to predict several future tokens, double as a built-in draft model at serving time. GLM-5.2’s report describes how the team optimized that MTP layer for serving by tuning draft cost, acceptance rate, and the loss it was trained with. Its serving design links the cost of generating drafts to the probability that the target accepts them.
The draft model moved inside
An MTP layer is a small extra transformer block sitting on top of the trunk. During training it predicts token from the hidden state that predicted , giving the model an additional prediction to learn from. At serving time you run it speculatively: the MTP layer drafts a chain of future tokens (GLM-5.2 uses 7 draft steps, with parameters shared across steps), and the target model verifies the whole chain in one forward pass, exactly the §18 protocol. Compared with a separate draft model, the MTP layer shares the trunk’s embeddings and hidden states, so its drafts start out well-aligned with the target distribution: think Medusa, but trained into the model from the start.
GLM-5.2’s team states the two objectives plainly: (1) make the MTP layer as cheap as possible as a draft model, and (2) make its drafts get accepted as often as possible. Every optimization below serves one of the two.
IndexShare and KVShare: cheaper drafts with consistent inputs
GLM-5.2 uses DeepSeek Sparse Attention (DSA)DeepSeek Sparse AttentionDeepSeek Sparse Attention (DSA): an attention variant where a lightweight indexer scores every past token with cheap dot products, a top-k selection keeps the most relevant ones, and full attention runs over only that subset.See in glossary →, which restricts full attention to selected past positions: a lightweight indexer scores past tokens and full attention runs over the top-k. IndexShareIndexShareSharing one sparse-attention indexer across a block of consecutive layers (or across speculative-decoding draft steps): the first computes which past tokens matter and the rest reuse its top-k selection, eliminating most indexer compute.See in glossary →, covered from the training side in the pre-training explainer, shares one indexer across every four backbone layers. The MTP layer gets the same treatment across time: the indexer runs only on the first draft step, and every following step reuses its top-k indices. The draft’s per-step cost drops accordingly (objective 1).
The surprise is that this same reuse also fixes an acceptance problem (objective 2). In multi-step drafting, GLM-5.1’s second draft step attended to a KV cacheKV cacheThe stored keys and values from all past tokens, so attention at step t only needs to compute Q for the new token.See in glossary → that was a mixture: keys and values for earlier positions came from the target model, but the newest entry came from the MTP layer’s own hidden state. During training, with teacher forcingteacher forcingDuring training, feeding the model the true previous tokens (not its own guesses) at every position, so all next-token predictions in a sequence can be learned in parallel.See in glossary →, which supplies the true preceding tokens during training, the MTP layer only ever saw target-model hidden states. So at inference it attended to a kind of history it had never seen in training: a classic train/inference mismatch, paid for in rejected drafts.
With IndexShare (plus the shared KV cache, KVShare), a later draft step reuses the first step’s indices and therefore attends only to positions whose KV came from the target model. What the draft sees at inference is now exactly what it saw in training. Training reuses the first draft step’s KV cache and indices the same way, closing the loop.
Rejection sampling and a loss that optimizes acceptance directly
The third and fourth improvements come from the Bebop paper (arXiv 2606.12370), which studied why MTP acceptance rates degrade, especially during reinforcement learning (RL)reinforcement learningLearning from trial and error: an agent takes actions and receives a single-number reward signal, with no labeled "right answer" for each step.See in glossary →, training that improves actions using reward scores, and traced the problem to fluctuations in output entropyentropyA measure of how spread-out (uncertain) a probability distribution is. In RL post-training, keeping entropy up preserves exploration and prevents premature collapse onto one answer.See in glossary →, a measure of how spread out the token probabilities are. Two fixes carry over to serving:
- Probabilistic rejection sampling instead of greedy drafting. Sample the draft tokens from the draft distribution and accept/reject them against the target distribution (the exact §18 protocol, which preserves the target distribution). Greedy drafts are brittle when the target distribution is flat; probabilistic drafting degrades gracefully.
- An end-to-end total-variation (TV)total variation distanceA measure of how much two probability distributions differ; for discrete outcomes, half the sum of the absolute differences between their probabilities.See in glossary → loss. Total variation measures the difference between two probability distributions as half the sum of their absolute probability differences. A lossloss functionA single number measuring how wrong the model's predictions are on a batch of data. Training works by adjusting the model to make this number smaller.See in glossary → is the numerical error minimized during training. Per-token cross-entropycross-entropy lossA classification loss that penalizes the model according to the negative log-probability it assigned to the correct answer.See in glossary →, a loss that penalizes low probability on the correct token, trains the MTP layer to be a good predictor, but what serving actually pays for is the multi-step acceptance rate of the whole drafted chain. The TV loss optimizes that quantity directly, end to end across draft steps.
The ablationablationAn experiment that removes or changes one part of a system to measure how much that part contributes to the result.See in glossary →, a comparison that isolates the effect of individual changes, measured as acceptance length on coding workloads (7 draft steps, GLM-5.1 backbone and data):
| Method | Acceptance length |
|---|---|
| Baseline | 4.56 |
| + IndexShare + KVShare | 5.10 |
| + Rejection sampling | 5.29 |
| + End-to-end TV loss | 5.47 (+20%) |
Kimi K3: turn the pre-trained head into an EAGLE draft
Kimi K3 starts from the same useful asset—a pre-trained MTP layer—but turns it into an EAGLE-3EAGLEA draft-model architecture that predicts feature vectors of the target model, achieving high acceptance rates.See in glossary →-style draft during post-training. The target model is frozen. Only the draft layer and a feature-fusion projection learn, taking low-, middle-, and high-level representations from K3’s Attention Residual blocks.
The initialization begins as the original MTP layer: the projection initially selects only the high-level feature it saw in pre-training, then gradually learns to use the lower-level signals. Training unrolls seven recurrent draft steps, feeding the draft its own prior outputs after the first step so training matches inference.
K3 also optimizes the quantity the server actually values. Instead of using Kullback–Leibler (KL) divergenceKL divergenceKullback–Leibler divergence — a measure of how far one probability distribution is from another. Used in post-training as a "leash" that keeps a model close to a reference policy.See in glossary →, a measure of how one probability distribution differs from another, as an indirect training target, its likelihood-based loss is the negative log of the draft–target overlap, —the exact per-token acceptance probability under lossless speculative sampling. GLM-5.2 and K3 arrive at the same broader principle with different mechanisms: a draft model should be trained for accepted tokens per target pass, not merely for generic next-token accuracy.
Serving a million tokens
The MTP work targets decode speed; GLM-5.2’s other serving headache is its headline feature. Extending context from 200K to 1M tokens shifts the bottleneck away from per-token FLOPs (which the sparse-attention work already cut) toward everything §12/§13 warned about: KV-cache capacity, kernels whose cost grows with context length, and CPU-side overhead in the serving engine. A 1M-token request’s KV cache is enormous even when its attention is cheap.
Z.ai reports optimizing along three fronts: finer-grained KV memory management and parallelization (building on LayerSplit) to reclaim usable cache space for ultra-long requests; coordinating long-context kernels with the KV-cache transfer pipeline so cache movement doesn’t stall prefill or decode; and CPU-side scheduling and runtime work to remove bubbles from the GPU pipeline. The result compounds with context length: normalized engine throughput vs GLM-5.1 goes from 1.03× at 32K to 6.97× at 1M, a regime where GLM-5.1 simply runs out of context.
Speculation becomes still trickier when accepting a token mutates a recurrent attention state that cannot simply be truncated. Kimi K3 addresses this by retaining the inputs needed to reconstruct the accepted state.