AI Fundamentals

Speculative Decoding: Why a Draft Model Can Speed Up a Larger LLM

AI Foundations #31 explains draft-and-verify generation, acceptance and rejection, why the target distribution can be preserved, and when speculative decoding actually reduces latency.

Approximately 6 min read · AI Foundations / Lesson 31

In AI Foundations #30, we looked at a memory problem: how an LLM server stores growing KV caches efficiently.

Now consider a latency problem.

Ordinary autoregressive generation produces tokens serially:

run target model
-> token 1
run target model
-> token 2
run target model
-> token 3

Speculative decoding asks:

What if something cheaper guesses several likely next tokens, and the large target model checks them together?

That is the core idea.

Draft and target

A basic system has two roles.

Draft: a smaller or cheaper predictor proposes several future tokens quickly.

Target: the model whose output distribution we actually want.

Suppose the text ends with:

The capital of France is

The draft might propose:

Paris . It is

Instead of asking the target model four times serially, the runtime can evaluate the proposed continuation in a more parallel verification pass.

If several draft tokens are accepted, the target advances multiple positions in one speculative round.

The target remains in charge

A common oversimplification is:

the small model writes and the big model approves.

The original speculative-sampling method is more precise.

The draft is only a proposal mechanism. The target probabilities determine acceptance, and a correction distribution is used when a proposal is rejected.

With the proper rule, sampling can preserve the target model’s distribution instead of silently switching to the draft model’s distribution.

A simplified round

Suppose the draft proposes:

d1 d2 d3 d4

Verification might produce:

d1 accepted
d2 accepted
d3 accepted
d4 rejected

The runtime samples a corrected token at the rejection point and starts another round.

The useful outcome was three accepted target tokens from one verification round instead of three separate serial target decode iterations.

Why verification can be cheaper

Transformers are good at processing multiple known token positions in parallel.

Autoregressive decode cannot fully exploit that because token N+1 depends on token N.

The draft temporarily breaks the dependency by guessing future tokens. The target can then score those known proposed positions together.

The bet is profitable when:

draft cost
+ verification cost
<
serial target cost for accepted progress

That makes acceptance rate central.

Acceptance rate controls the economics

If the draft closely matches the target, many tokens survive each round.

If the draft disagrees early, most speculative work is discarded.

So speculative decoding is not automatically faster. Its value depends on:

The useful metric is accepted target progress per verification round, not simply the number of draft tokens proposed.

The smallest draft is not always best

A tiny draft can be very fast but disagree frequently with the target.

A larger draft can agree more often but consume too much compute.

So there is a trade-off:

tiny draft
-> low proposal cost
-> possibly low acceptance

larger draft
-> higher acceptance
-> higher proposal cost

The best point depends on the model pair and workload.

The proposer does not have to be another full LLM

Current vLLM documentation lists several speculative methods, including:

The proposer changes, but the high-level structure remains:

propose
-> verify
-> accept useful prefix
-> correct at first failure
-> repeat

Some models include heads or modules trained to predict several future tokens.

Those predictions can become speculative candidates without loading a completely separate draft model.

That can reduce proposer overhead and improve alignment with the target, although the runtime still needs an efficient verification path.

What speculation is trying to reduce

The main target is serial inter-token latency.

vLLM documentation frames speculative decoding as particularly useful for reducing inter-token latency in suitable medium-to-low-QPS, memory-bound workloads.

At high server load, the result can change. Normal continuous batching may already keep the accelerator busy, while the draft consumes capacity that could have served additional target requests.

Latency and maximum aggregate throughput are different objectives.

Why high concurrency can shrink the advantage

At low concurrency, one large model decoding one token can underuse the GPU.

Speculation widens the verification work.

At high concurrency, the server already has token positions from many requests to batch together. Extra draft compute can become overhead rather than useful parallelism.

A method can therefore improve single-user latency while failing to improve saturated-server throughput.

Memory is part of the trade-off

A separate draft model needs weights and runtime state.

That memory could otherwise support a larger target model, more KV cache, longer contexts, larger batches or more replicas.

So the correct comparison is whole-system performance, not draft-model speed in isolation.

Speculation length has a sweet spot

Proposing very few tokens leaves little opportunity to skip target iterations.

Proposing many tokens increases theoretical progress, but an early rejection wastes more draft work.

If a proposal had an independent per-token acceptance probability p, the chance that k tokens all survive would roughly shrink like:

p raised to k

Real acceptance is not independent, but the intuition explains why longer speculation is not always better.

Greedy decoding is the simplest mental model

For deterministic greedy output, the idea is intuitive:

draft guesses
target checks
keep matching prefix

Sampling needs the proper acceptance and correction rule because several tokens can be valid probabilistic outcomes.

That distribution-preservation property is what makes speculative decoding more than a heuristic shortcut.

Benchmark it correctly

Keep the target model and request set fixed.

Compare:

baseline target only
vs
target plus speculation

Record:

A speedup without acceptance data is hard to interpret.

Main failure modes

Speculation can disappoint when:

The draft disagrees too often. Too many proposals are discarded.

The draft is too expensive. Saved target iterations are replaced by nearly equal proposal cost.

Verification is inefficient. Multi-token verification does not exploit enough parallelism.

The server is already saturated. Extra speculative work competes with normal batching.

Memory pressure reduces batch capacity. The proposer consumes memory the target needed for KV cache or concurrency.

Mental model to keep

Speculative decoding does not make the target model smaller.

It tries to reduce the number of serial target decode iterations required for the same accepted progress.

The core equation is:

cheap accurate guesses
+ efficient parallel verification
= fewer expensive serial target steps

That is why speculative decoding can be a major latency optimization in one workload and nearly useless in another.

Sources and further reading

Continue reading