Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 7 min read

Speculative Decoding Made My vLLM Server Slower: The Acceptance Math

I added a draft model to my vLLM server, sent one curl request, and watched the tokens fly. Then I ran the load test with 64 concurrent users and total throughput went down. Same GPU, same target model, same prompts. The

I added a draft model to my vLLM server, sent one curl request, and watched the tokens fly. Then I ran the load test with 64 concurrent users and total throughput went down. Same GPU, same target model, same prompts. The only change was one flag that every blog post says is a free speedup.

Speculative decoding is not free. It is a bet, and the odds depend on two numbers most people never look at: the acceptance rate and how busy your GPU already is. This post is the math I wish I had done before flipping the flag.

TL;DR

  • Speculative decoding lets a small draft model guess the next few tokens, then the big model verifies all of them in a single forward pass. Accepted guesses are free tokens.
  • Expected tokens per big-model step is (1 - ฮฑ^(ฮณ+1)) / (1 - ฮฑ), where ฮฑ is the acceptance rate and ฮณ is how many tokens you draft. At ฮฑ = 0.8, ฮณ = 4 that is 3.36 tokens per step.
  • That speedup only exists while decoding is memory-bandwidth bound (small batches). At large batch sizes the GPU becomes compute bound, verifying 5 tokens costs roughly 5x, and rejected drafts are pure waste. The same config can drop to about 0.65x.
  • Higher temperature and off-domain prompts lower ฮฑ. Measure it from vLLM's /metrics before you trust any speedup.
  • Fix: keep num_speculative_tokens small (2 to 4), disable speculation above a batch size threshold, and use n-gram (prompt lookup) drafting for edit-heavy workloads.

What is speculative decoding actually doing?

Speculative decoding splits generation into a cheap guesser and an expensive judge. A small draft model (say Llama 3.2 1B) proposes ฮณ tokens one at a time. The target model (say Llama 3.1 70B) then runs one forward pass over all ฮณ proposed positions and checks each guess.

The check walks left to right. The first guess that fails is replaced by a token sampled from a corrected distribution, and every guess after it is thrown away. So one target step always yields at least 1 token and at most ฮณ + 1.

Two details matter for production:

  1. The output distribution is unchanged. With proper rejection sampling, the tokens you get follow exactly the target model's distribution. Speculative decoding changes speed, not quality. If your quality changed, something else is wrong.
  2. The draft must share the target's tokenizer. The verification compares token IDs. A draft model with a different vocabulary is not a draft model, it is a random number generator.

How many tokens does speculative decoding save per step?

The expected number of tokens per target forward pass, assuming each draft token is accepted independently with probability ฮฑ, is:

E[tokens per step] = (1 - ฮฑ^(ฮณ+1)) / (1 - ฮฑ)

The speedup also has to pay for the draft. If one draft forward pass costs c times a target pass, and verification costs about one target pass, the speedup is:

speedup = E[tokens per step] / (1 + ฮณยทc)

Here is that formula with ฮณ = 4 and c = 0.05 (an illustrative value; draft passes are often more expensive than the parameter ratio suggests because of kernel launch overhead):

Acceptance rate ฮฑ Tokens per step Speedup at batch 1
0.9 4.10 3.41x
0.8 3.36 2.80x
0.6 2.31 1.92x
0.4 1.65 1.37x
0.2 1.25 1.04x

Halving ฮฑ from 0.8 to 0.4 wipes out most of the speedup. And drafting more tokens barely helps a weak draft: at ฮฑ = 0.4, going from ฮณ = 4 to ฮณ = 8 moves tokens per step from 1.65 to 1.67 while the cost per step goes from 1.2 to 1.4 target passes. Speedup drops to about 1.19x.

That table is the optimistic case. My load test lived somewhere else.

Why does speculative decoding get slower under load?

Speculative decoding gets slower under load because its whole trick depends on verification being nearly free, and that is only true when the GPU is waiting on memory, not on math.

At batch size 1, decoding is memory-bandwidth bound. Every token requires streaming all the model weights out of HBM, and the tensor cores mostly sit idle. Run the numbers with public H100 specs: roughly 990 TFLOPS of dense BF16 against roughly 3.35 TB/s of memory bandwidth. That is about 295 FLOPs per byte before compute becomes the bottleneck.

In BF16, each weight is 2 bytes and contributes 2 FLOPs per token (one multiply, one add). So one token per weight read is about 1 FLOP per byte. You can push on the order of a few hundred tokens through each weight read before the math catches up with the memory. Below that, extra tokens in the same forward pass are almost free. That is the gap speculative decoding exploits.

Now do the math for my load test: 64 concurrent sequences, each verifying 4 draft tokens plus 1 bonus position. That is 320 token positions per forward pass. Already past the theoretical ridge point, and real kernels hit their compute ceiling earlier than the spec sheet suggests, and KV cache reads eat bandwidth on top.

Once you are compute bound, verifying ฮณ + 1 positions costs roughly ฮณ + 1 times as much as decoding one. Plug that into the formula as the worst case:

speedup = E[tokens per step] / ((ฮณ + 1) + ฮณยทc)
Acceptance rate ฮฑ Speedup when compute bound (ฮณ = 4)
0.9 0.79x
0.8 0.65x
0.6 0.44x
0.4 0.32x

Every row is below 1. Since E can never exceed ฮณ + 1, a fully compute-bound server can never win with speculation. Every rejected draft token is FLOPs you paid for and threw away, and those FLOPs came out of other users' requests.

Reality sits between the two tables, set by batch size, model size, and hardware. My single curl and my load test were both correct.

Why does temperature lower the acceptance rate?

Temperature lowers the acceptance rate because acceptance depends on how much the draft and target distributions overlap, and sampling spreads probability mass over more tokens that the two models disagree on.

Under rejection sampling, the probability that a draft token is accepted at a given position is ฮฃ min(p(x), q(x)) over the vocabulary, where p is the target distribution and q is the draft distribution. With greedy decoding this reduces to a simple question: do both models have the same argmax? On boilerplate code, they usually do.

Turn temperature up and both distributions flatten into long tails that rarely match. Domain drift does the same: a 1B model that predicts JSON keys perfectly can be useless on legal prose.

So "speculative decoding gave a 2.5x speedup" is a claim about one model pair, one temperature, one domain, and one batch size. Change any of them and the number changes.

How do I check if speculative decoding is helping my vLLM server?

Measure the acceptance rate on your real traffic, then plug it into the formula for your actual batch size. vLLM exposes speculative decoding counters on its Prometheus endpoint:

curl -s localhost:8000/metrics | grep spec_decode

On recent versions you get counters for the number of drafts, drafted tokens, and accepted tokens (exact metric names vary by version, so grep instead of hardcoding). Two numbers fall out:

  • accepted / drafted is your per-token acceptance rate.
  • 1 + accepted / drafts is your mean tokens per target step.

Then estimate before you benchmark:

def spec_speedup(alpha, gamma, c, verify_cost=1.0):
    """verify_cost: ~1.0 when memory bound, up to gamma+1 when compute bound."""
    expected = (1 - alpha ** (gamma + 1)) / (1 - alpha)
    return expected / (verify_cost + gamma * c)

for verify in (1.0, 2.0, 3.0, 5.0):
    print(verify, round(spec_speedup(0.7, 4, 0.05, verify), 2))

With ฮฑ = 0.7, that prints roughly 2.31, 1.26, 0.87, and 0.53. Your break-even point is where verify_cost crosses about 2.6. A load test at your real concurrency tells you which side you are on. A single curl tells you nothing.

What should I actually configure?

Treat speculation as a low-traffic optimization, not a default. Here is what I changed:

1. Keep the draft short. num_speculative_tokens of 2 to 4. Long drafts only pay off when ฮฑ is above about 0.8, and long drafts hurt the most once you are compute bound.

vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --speculative-config '{"model": "meta-llama/Llama-3.2-1B-Instruct", "num_speculative_tokens": 3}'

2. Turn it off when the batch fills up. Some vLLM versions support disabling speculation above a batch size threshold (look for disable_by_batch_size in the speculative config for your version). If yours does not, route latency-sensitive low-concurrency traffic to a speculative deployment and bulk traffic to a plain one.

3. Try n-gram drafting for edit workloads. If your outputs copy large spans of the input (code edits, document rewrites, structured extraction), prompt lookup drafting finds candidate tokens in the prompt itself. No draft model, near-zero draft cost, and high acceptance on copy-heavy text.

--speculative-config '{"method": "ngram", "num_speculative_tokens": 4, "prompt_lookup_max": 4}'

4. Measure per workload. One average acceptance rate across chat and batch traffic hides the workload where speculation is losing money.

So does speculative decoding make vLLM faster?

Speculative decoding makes vLLM faster only when the GPU is memory-bandwidth bound and the draft model agrees with the target often enough. The expected tokens per step is (1 - ฮฑ^(ฮณ+1)) / (1 - ฮฑ), which gives about 2.8x at batch size 1 with an 80% acceptance rate and 4 draft tokens. At high concurrency the GPU becomes compute bound, verifying draft tokens stops being free, and the same configuration can fall to around 0.65x. Measure your acceptance rate from /metrics, keep drafts short, disable speculation at large batch sizes, and benchmark at your real concurrency, not with a single request.

Written by the developer behind Preterview, an interview prep platform.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.