Section 04

Structured outputs and faster decoding

Constraints, drafts, and verification

The handbook assistant now returns a machine-readable object: a policy name, a short answer, and a list of supporting sections. Its application must parse that object before it can display anything. Meanwhile, users want long answers to arrive faster after the first token.

These requirements point to two different runtime mechanisms. Structured-output decoding restricts which continuations are legal. Speculative decoding proposes continuations cheaply and verifies them with the target model. Both change the generation loop, but they obtain their benefits for different reasons and offer different guarantees.

A grammar restricts the next choice

Suppose an application requires an object with fields policy and answer. At some positions, a quote or colon is required. At others, the model may choose among many strings. A constrained decoder tracks where generation is within the allowed structure and masks tokens that would make it invalid.

The mask operates on tokens, while the structure is usually described in characters or higher-level syntax. A single token may contain multiple characters, cross a boundary, or encode a piece of punctuation together with a space. The runtime needs a grammar implementation that accounts for the tokenizer’s vocabulary rather than pretending that every character is a token.

SGLang’s current structured-output documentation describes interfaces for JSON schemas, regular expressions, and grammar constraints, with configurable backends. The request format and supported schema features belong to that backend and version. Structured outputs documentation

Our assistant can benefit even when constraints do not increase raw tokens per second. If an unconstrained response would require a repair call, preventing the syntax error avoids an entire additional inference request. The relevant application metric might be time to a usable result, not simply the number of generated tokens.

The original compressed-state-machine idea

The SGLang paper explored a further opportunity: some grammar paths contain a fixed sequence with no choice along the way. A compressed finite-state machinecompressed finite-state machineA grammar representation that combines deterministic paths; the original SGLang design used it to organize forced output tokens more efficiently.See in glossary → combines such deterministic stretches, allowing their processing to be organized more efficiently than repeatedly making a one-token decision. SGLang paper, §4

Consider a toy format that begins with the literal text {"policy":. Once the decoder has committed to that structure, some following characters may be forced. The model still needs the state resulting from processing those tokens. Knowing the required text can allow several positions to be handled together; it does not permit omitting the KV updates that later attention depends on.

This is a historical design explanation. It should not be read as a promise that every modern SGLang grammar backend automatically compresses every deterministic path in the same way. The supported backend, tokenizer interactions, and constrained sampling behavior matter. The paper itself discusses tokenization complications and probability distortion in its appendices.

Syntax is only one layer of correctness

A valid JSON object can still contain a false claim. A permitted policy name can be paired with an irrelevant citation. A syntactically valid tool call can request the wrong operation. Grammar enforcement controls the shape of a response within the supported constraint language; it does not verify the factual reasoning behind that response.

For the handbook application, validate that cited section IDs exist and that required fields are present. Evaluate whether the answer is supported by the document. If the task requires a semantic verifier, account for its inference cost separately. Structured output makes integration easier, but it does not eliminate the rest of the application’s correctness requirements.

A constraint can also change the distribution of outputs. Masking forbidden tokens and renormalizing the rest defines a constrained generation process. Comparing it with unconstrained sampling is not an apples-to-apples speed comparison unless you also explain that behavioral difference. This becomes especially important when constraints and speculative verification interact.

Predict cheaply, verify in bulk

The inference companion’s speculative-decoding chapter introduced a draft model that proposes several tokens. The target model then evaluates those candidate positions in a more parallel operation than ordinary one-step-at-a-time generation.

SGLang supports multiple speculative methods, including EAGLE-family approaches and model-specific multiple-token prediction paths. EAGLE is based on predicting target-model features rather than simply treating any small language model as an interchangeable draft. Configuration and compatibility depend on the target and draft models. SGLang speculative decoding, EAGLE paper

The useful quantity is not the number of drafted tokens. It is the number of tokens actually emitted per unit of total work. A candidate branch that is mostly rejected can increase computation without much output progress. Draft memory, verification shapes, batch size, and host coordination all contribute to cost.

Here is a deliberately simple example. Ordinary decoding emits one token every 20 milliseconds. A speculative cycle spends 8 milliseconds drafting, 24 milliseconds verifying, and 2 milliseconds on other overhead. Its total is 34 milliseconds. If it emits an average of three tokens, the effective interval is about 11.3 milliseconds per emitted token. If it emits only one, it is slower than the ordinary decoder.

The break-even condition in this toy model is:

draft time+verify time+other timemean emitted tokens per cycle<ordinary time per token.\frac{\text{draft time} + \text{verify time} + \text{other time}} {\text{mean emitted tokens per cycle}} < \text{ordinary time per token}.

Count the emitted tokens consistently. Some algorithms can emit a correction or bonus token beyond accepted draft tokens. An “acceptance length” reported by one implementation may include that token while another metric does not. Read the metric definition before inserting it into a speedup calculation.

Exactness depends on verification

For greedy decoding, a verifier can accept a matching prefix of the target’s greedy choices, subject to implementation details and numerical behavior. For stochastic sampling, a suitable acceptance and correction rule is needed to preserve the target distribution. Blindly accepting tokens because they look plausible does not provide that guarantee. The original speculative sampling papers derive the relevant procedure. Fast Inference from Transformers via Speculative Decoding

Even when an algorithm preserves the distribution mathematically, two runs need not produce identical strings with the same random seed. Kernel numerics, scheduling, and the consumption of random numbers can differ. Distinguish distributional correctness from bitwise determinism and from matching a particular sampled transcript.

If a grammar is active, the verifier and draft process must respect the intended constrained target distribution. Compatibility support is something to check in the chosen release, not infer from the fact that two features appear separately in a feature list.

Speculation changes the serving system too

A draft model consumes memory that could otherwise hold active requests or reusable prefixes. Verification batches may use different kernels and shapes. More output per target pass can change CPU scheduling pressure. An optimization that helps a single user at low concurrency can reduce the number of users the server accommodates under memory pressure.

The SGLang team’s MTP discussion reports integration with prefill–decode disaggregation and expert parallelism for supported configurations. That is useful evidence of implemented combinations; its measurements apply to the tested models and hardware. It does not establish that every speculative method composes with every transfer path. SGLang MTP implementation discussion

For our assistant, test four cases: ordinary decoding, constrained decoding, speculative decoding, and the supported combination. Keep prompts, output requirements, arrival rates, and hardware fixed. Measure usable answers per second as well as output token rate. Report failures and truncations rather than removing them from the denominator.

A service serving short structured answers may gain more from avoiding repair calls than from speculative acceleration. A service generating long unconstrained reports may have the opposite profile. Both mechanisms are part of the runtime’s toolkit, but their value comes from the application’s actual output behavior.

Sources and further reading