AI Fundamentals

Prefix Caching and KV Reuse: Why Repeated Prompts Can Be Cheaper

AI Foundations #29 explains prefix caching, KV reuse, exact token matching, cache hits, eviction, prompt-template effects, and why reuse reduces prefill work without caching the answer.

Approximately 5 min read · AI Foundations / Lesson 29

In AI Foundations #28, we looked at how an LLM server decides whose work runs next.

Now consider a different question:

What if many requests begin with exactly the same prompt?

That can happen with the same system prompt, long document, tool definitions, conversation history or few-shot examples.

A serving engine may be able to reuse computation from that shared beginning. That is prefix caching.

First remember what prefill does

When a prompt enters a transformer, the model processes the input tokens before it can generate the next token.

For:

t1 t2 t3 ... tN

the model computes intermediate attention state. That work is prefill.

The resulting key/value state is stored in KV cache so later decode steps can attend to earlier tokens without recomputing the whole prefix.

Repeated prefixes create repeated work

Imagine every request starts with a 6,000-token policy document.

User A sends:

[shared 6000-token prefix]
How do I reset my account?

User B sends:

[shared 6000-token prefix]
How do I change my address?

Without reuse, the server processes those 6,000 tokens twice.

With prefix caching, it may process the shared prefix once and reuse its KV state for the second request.

This is not answer caching

Prefix caching does not mean:

same question -> return an old answer

It means:

same input prefix
-> reuse already-computed model state
-> prefill only the new suffix
-> generate normally

Decode still happens. Sampling can still produce different output.

Matching happens at token level

Text that looks identical is not enough if tokenization or serialization differs.

A cacheable prefix is closer to:

same token IDs
in the same order
for the same model/runtime context

Small differences can break reuse: whitespace, chat templates, tool ordering, tokenizer changes or an early system-message change.

Prompt construction becomes part of serving performance.

A simple example

Suppose a cached prefix covers tokens 1 through 4,000 and a new request contains 4,300 tokens.

If the first 4,000 match exactly:

reuse KV for 1..4000
+ compute 4001..4300
= complete prompt state

The exact block boundaries depend on the runtime.

Why block-based KV memory helps

Serving engines often manage KV state in blocks rather than one giant allocation per request.

That makes reuse practical. A later request can reference matching blocks and allocate new blocks only for its suffix.

This connects prefix caching to paged KV-memory designs.

SGLang pushes reuse into prompt structure

SGLang’s RadixAttention work organizes reusable state around shared prefixes.

A useful mental model is:

system prompt
|
+-- shared document
    |
    +-- question A
    |
    +-- question B

Requests that share the upper path can reuse the associated state.

The first request still pays

The cache has to exist before anyone can hit it.

A workload with no repeated prefixes may see little benefit.

A reuse-heavy workload can look more like:

one expensive first prefill
+ many cheaper matching requests

Average visible prompt length no longer tells you the real amount of prefill compute.

Cache hits can reduce TTFT

If much of a prompt is already cached, less new prefill work is required before decode.

That can reduce time to first token.

But the gain depends on match length, cache residency, queueing, scheduler policy, backend speed and uncached suffix length.

Caching removes work. It does not remove contention.

Reuse consumes memory

Reusable state must remain resident somewhere.

Keeping more prefixes can improve hit rate, but it occupies cache capacity.

The trade is:

more cached prefixes
-> more possible hits
-> more memory occupied

When space is needed, old entries may be evicted.

Prompt templates can destroy reuse

Two requests may contain the same tools but serialize them in a different order.

Humans see the same meaning. The tokenizer sees a different token stream.

That can turn a large hit into almost no hit.

Stable serialization therefore matters for applications that depend on prefix reuse.

Multi-turn chat is naturally prefix-shaped

Conversation history grows:

turn 1
-> turn 1 + turn 2
-> turn 1 + turn 2 + turn 3

Each later request contains an earlier prefix, so a serving engine may reuse history state rather than rebuilding from token zero.

Reuse changes capacity planning

Without caching, a rough model might be:

requests per second
x average prompt tokens

With caching, separate:

total prompt tokens
cached prompt tokens
new prompt tokens
decode tokens

Two workloads with the same visible prompt length can have very different compute cost.

What prefix caching cannot fix

It does not fix slow decode, insufficient GPU memory, an overloaded scheduler, a long uncached suffix or a prompt that changes near the beginning every time.

It specifically attacks repeated prefix computation.

The cumulative serving map

The sequence now becomes:

request arrives
-> scheduler admits work
-> prefix lookup finds reusable state
-> uncached tokens run through prefill
-> KV state is stored
-> decode generates
-> scheduler interleaves requests

The main takeaway

Prefix caching is computation reuse, not answer caching.

The server stores model state for an exact token prefix so a later request can continue from an already-computed point.

That makes prompt stability, KV memory, eviction policy and scheduling part of the same serving story.

Sources and further reading

Continue reading