Dev.to AI 🤖 Ai 👁 0 📖 8 min read

How Structured Outputs Actually Work: Grammar-Constrained Decoding and Logit Masking Under the Hood

When OpenAI launched response_format={"type": "json_schema", "strict": true} and open-source inference engines like vLLM, SGLang, and llama.cpp introduced grammar-constrained decoding, developer workflows changed overnig

When OpenAI launched response_format={"type": "json_schema", "strict": true} and open-source inference engines like vLLM, SGLang, and llama.cpp introduced grammar-constrained decoding, developer workflows changed overnight.

Before this, getting an LLM to reliably return valid JSON required prayers, complex system prompts ("You are a JSON generator. Do NOT include markdown code blocks"), and brittle retry loops wrapped around json.loads(). Even with fine-tuned models, a rogue comma or unescaped quote would inevitably crash downstream production pipelines.

Today, structured outputs are guaranteed to be 100% syntactically valid according to your exact JSON schema.

Most developers assume one of two things is happening behind the scenes:

  1. The model was fine-tuned specifically to follow JSON schemas.
  2. The API server runs a hidden while loop, catching syntax errors and asking the model to fix its mistakes.

Both assumptions are completely wrong.

In reality, the model never gets the chance to make a syntax error. The inference engine physically alters the model's output probabilities at every single token generation step.

Here is the exact systems mechanics of how grammar-constrained decoding and logit masking work under the hood.

1. The Normal Autoregressive Decoding Loop

To understand grammar constraints, we first need to look at how an LLM produces a single token during unconstrained text generation.

At decoding step $t$, the transformer processes the input context and KV cache, outputting a vector of unnormalized raw scores called logits across its entire vocabulary $V$:

Logits vector z: [z_1, z_2, z_3, ..., z_|V|]
Vocabulary size |V|: ~128,000 tokens (for Llama 3 / GPT-4)

In standard greedy or temperature-based sampling, these logits are passed through a Softmax function to convert them into a probability distribution:

$$P(w_i) = \frac{e^{z_i / T}}{\sum_{j=1}^{|V|} e^{z_j / T}}$$

The sampler draws a token $w_t$ based on these probabilities.

In unconstrained decoding, every token in the vocabulary has a non-zero probability. Even if there is only a 0.0001% chance of sampling a closing brace } right after a key name "username":, over billions of generated tokens, that failure will happen thousands of times in production.

2. Step 1: Compiling JSON Schema into a Formal Grammar State Machine

When you submit a JSON Schema or Pydantic model to an inference engine (such as vLLM's xgrammar, SGLang's outlines, or llama.cpp's GBNF engine), the engine does not feed the schema to the model weights. Instead, it compiles the schema into a formal grammar state machine on the host CPU or GPU.

[ JSON Schema / Pydantic Model ]
               │
               ▼
   [ Schema Compiler Engine ]
               │
               ▼
[ Pushdown Automaton (PDA) / DFA ]
  - States: Start, InKey, ColonExpected, InValue, CommaOrClose, etc.
  - Stack: Tracks nested object '{' and array '[' depth.

JSON is a Context-Free Grammar (CFG) because it supports arbitrarily nested arrays and objects ({"a": {"b": [1, 2]}}). A standard Deterministic Finite Automaton (DFA) cannot track arbitrary nesting depth because it has finite memory. Therefore, the engine compiles the schema into a Pushdown Automaton (PDA): a finite state machine equipped with a stack.

Let us trace a simple schema requiring an object with a string name and an integer age:

  1. State 0 (Start): Only valid character is {.
  2. State 1 (Expect Key): Only valid character sequence is "name".
  3. State 2 (Expect Colon): Only valid character sequence is :.
  4. State 3 (Expect String Value): Valid characters are ", followed by any valid JSON string characters, terminated by ".
  5. State 4 (Expect Delimiter): Only valid character is ,.
  6. State 5 (Expect Next Key): Only valid character sequence is "age".
  7. State 6 (Expect Colon): Only valid character sequence is :.
  8. State 7 (Expect Integer): Only valid characters are digits 0-9.
  9. State 8 (Expect Close): Only valid character is }.
  10. State 9 (End): Only valid token is the End-of-Sequence token (<|endoftext|>).

At every character in the output, the PDA knows precisely which next characters are legally allowed.

3. Step 2: The Tokenizer Vocabulary Problem

Here is the central challenge that makes grammar-constrained decoding difficult: Grammars operate on characters, but LLMs operate on tokens.

Modern tokenizers (like Byte-Pair Encoding / BPE) group multi-character sequences into single token IDs. In Llama 3's 128k vocabulary:

  • Token 1234 might represent "hello"
  • Token 5678 might represent ": true,"
  • Token 9012 might represent " 42 "
  • Token 3456 might represent ": {"

If the grammar is currently in State 2 (expecting a colon : followed by a space and a boolean), which of the 128,000 token IDs in the vocabulary are allowed?

To solve this without re-evaluating 128,000 strings on every step, inference engines pre-build a Vocabulary Prefix Trie.

                [Root]
               /      \
             """     "{"
             /   \      \
          "age"  "name"  ""user""
           /        \
       "": "      "": "

The Vocabulary Trie indexes every token in the model's vocabulary byte-by-byte.

When the grammar is at state $S_t$, the engine traverses the Trie against the current state of the Pushdown Automaton. A token ID is marked valid if and only if every character in the token's byte string represents a valid path in the grammar state machine.

4. Step 3: Runtime Logit Masking (The Mathematical Filter)

Once the engine determines the subset of valid tokens $V_{\text{valid}} \subseteq V$ for the current state $S_t$, it constructs a Logit Mask Vector $M \in \mathbb{R}^{|V|}$:

$$M_i = \begin{cases} 0 & \text{if } i \in V_{\text{valid}} \ -\infty & \text{if } i \notin V_{\text{valid}} \end{cases}$$

This mask is added directly to the raw logits vector $z$ output by the transformer forward pass before Softmax is computed:

$$\tilde{z} = z + M$$

$$\tilde{z}i = \begin{cases} z_i & \text{if } i \in V{\text{valid}} \ -\infty & \text{if } i \notin V_{\text{valid}} \end{cases}$$

Now, let us observe what happens when we calculate the Softmax probability for an invalid token $k \notin V_{\text{valid}}$:

$$P(w_k) = \frac{e^{\tilde{z}k}}{\sum{j=1}^{|V|} e^{\tilde{z}j}} = \frac{e^{-\infty}}{\sum{j=1}^{|V|} e^{\tilde{z}j}} = \frac{0}{\sum{j=1}^{|V|} e^{\tilde{z}_j}} = 0$$

Because $e^{-\infty} = 0$, the probability of every invalid token becomes mathematically zero.

When the sampler selects the next token, it is physically impossible to choose anything that violates the JSON schema. The model's learned knowledge and attention weights decide which valid token to pick, but the grammar defines the sandbox.

[ Transformer Forward Pass ] ──► Raw Logits: [ 4.2,  1.1,  8.7,  -0.5, ... ] (128k items)
                                                    │
[ Grammar State + Vocab Trie ] ─► Mask Vector: [ 0.0, -inf,  0.0,  -inf, ... ]
                                                    │
                                                    ▼
                                Masked Logits: [ 4.2, -inf,  8.7,  -inf, ... ]
                                                    │
                                                    ▼
                                           [ Softmax Sampling ]
                                                    │
                                                    ▼
                                         Guaranteed Valid Token

5. Step 4: Deterministic Jump Decoding (Token Fast-Forwarding)

Consider what happens when the model generates the key "is_admin".

According to our schema, the key name is fixed. The colon : is fixed. If the field is a boolean, the only allowed tokens are true or false.

Why execute heavy matrix multiplications across 70 billion parameters on an expensive GPU just to predict the characters ": " or the closing quote of a known key?

Modern constrained decoding engines implement Jump Decoding (also called Token Infill or Speculative Schema Fast-Forwarding):

  1. At state $S_t$, the engine evaluates the set of valid next tokens $V_{\text{valid}}$.
  2. If $|V_{\text{valid}}| = 1$ (there is exactly one legal token), or if the schema specifies a literal string constant, the engine skips the transformer forward pass completely.
  3. It directly appends the known token to the output sequence, updates the KV cache, advances the grammar state machine, and continues to the next decision point.
Output: { "status": "active" }
          ▲       ▲▲        ▲▲
          │       ││        ││
Forward Pass?   NO (Fixed)  NO (Schema Key)  YES (Value Choice)  NO (Closing)

In structured extraction tasks where 40% to 60% of the output tokens are predictable schema scaffolding (keys, brackets, whitespace, punctuation), jump decoding can increase overall generation throughput by 2x to 3x compared to unconstrained decoding.

6. The Engineering Bottleneck: Masking Latency vs GPU Throughput

While logit masking sounds simple in theory, naive implementations create severe inference bottlenecks.

Consider high-throughput serving:

  • A batch of 64 requests generating tokens on an NVIDIA H100.
  • Vocabulary size: 128,000 tokens.
  • Each decode step takes ~8ms on GPU.
  • If evaluating the PDA and updating 128,000 logits in Python on the CPU takes 10ms per step, the GPU spends more time waiting for the CPU mask than doing tensor arithmetic!

How Modern Engines Solve This (vLLM xgrammar & SGLang):

  1. Pre-allocated Bitset Lookup Tables: Rather than computing masks on the fly, static states (such as keywords, boolean branches, and punctuation) are precompiled into bitsets where each bit corresponds to a token ID in $V$.
  2. Custom CUDA Logit-Masking Kernels: The bitset is copied to GPU memory once. A dedicated CUDA/Triton kernel applies the $-\infty$ mask in parallel across all 128k logits in less than 50 microseconds (0.05ms).
  3. Adaptive Token Trie Pruning: High-frequency prefixes are cached in L1/L2 CPU cache, minimizing pointer chasing during grammar transitions.

7. Practical Gotchas Developers Must Know

Even though structured outputs guarantee syntax validity, they introduce real engineering trade-offs:

1. Model Distribution Distortion

When you force an LLM to follow a strict schema, you are truncating its natural probability distribution. If your schema forces the model to output a specific field order that conflicts with how it reasons (for example, asking for "final_answer" before "explanation"), the model cannot perform step-by-step reasoning. Always put reasoning or thought fields before categorical conclusions in your schemas.

2. The Unbounded String Trap

A schema with {"type": "string"} allows any sequence of characters inside quotes. The grammar engine will ensure the string opens and closes with ", but it cannot prevent the model from generating 4,000 tokens of repetitive text inside that string. Always use length constraints, enums, or regex patterns if you need bounded string outputs.

3. Schema Compilation Overhead

Compiling complex nested schemas with hundreds of union types (anyOf, oneOf) into a Pushdown Automaton can take 50ms to 500ms on the first request. Production systems solve this by precompiling and warming schema state machines at server startup rather than compiling per request.

Summary: The 4-Step Mental Model

Whenever you use Structured Outputs in production, remember what is really happening:

  1. Schema to PDA: Your JSON schema is compiled into a Pushdown Automaton state machine with an explicit stack for tracking object/array nesting.
  2. Vocab Trie Traversal: The tokenizer's vocabulary is matched against the automaton to find the exact subset of valid token IDs.
  3. Logit Masking: Invalid token positions are overwritten with $-\infty$ before Softmax, reducing their sampling probability to exactly 0%.
  4. Jump Decoding: Deterministic schema scaffolding (keys, colons, brackets) bypasses GPU computation entirely and injects directly into the KV cache.

Guaranteed JSON from LLMs is not magic prompt engineering; it is deterministic formal language theory controlling probabilistic token sampling.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.