AI Fundamentals

Context and KV Cache: How an LLM Reuses Earlier Tokens

AI Foundations #17 explains context windows and KV cache: what the model can see, what keys and values are cached, why decoding gets faster, and why long chats consume more memory.

Approximately 11 min read · AI Foundations / Lesson 17

Yesterday’s lesson ended with sampling.

We had reached this loop:

model reads current tokens
-> produces logits
-> sampler chooses one token
-> append that token
-> repeat

That raised a new question for me.

Suppose the prompt already contains 5,000 tokens and the model generates one more token.

Does it really have to recalculate everything about those same 5,000 earlier tokens again?

Then after token 5,001, does it recalculate the first 5,000 again?

And again at 5,002?

That sounds incredibly wasteful.

The answer is that transformer inference normally uses a KV cache so it can reuse important attention information from earlier tokens.

To understand that, I first need to separate two ideas that I used to mix together:

context window
KV cache

They are related, but they are not the same thing.

Context is the token history the model can use

The context is the sequence of tokens available to the model for the current prediction.

If I write:

The company reported revenue of $10 million last year.
This year revenue increased by 20%.
What is this year's revenue?

all of those prompt tokens are part of the context.

If the model has already generated:

This year's revenue is

those generated tokens become part of the context too.

So generation keeps growing the sequence:

prompt tokens
+ generated token 1
+ generated token 2
+ generated token 3
+ ...

The model uses that growing history when predicting the next token.

The context window is the maximum usable history

A model cannot usually keep an unlimited number of tokens active.

It has some maximum context length, often called the context window.

A simplified example:

maximum context = 8,192 tokens

If the prompt uses:

7,000 tokens

then only about:

1,192 tokens

remain before the total sequence reaches 8,192.

That total normally includes both input and generated tokens.

So:

prompt length + generated length <= usable context limit

The exact behavior when the limit is reached depends on the runtime and application. A system may reject the request, truncate older tokens, shorten requested output, or use another context-management strategy.

But the basic idea is simple:

the context window describes how much token history can participate in the model’s current computation.

Context is not long-term memory

This distinction helped me a lot.

If I tell a model my favorite company inside one conversation, it can use that information because the relevant tokens remain in context.

That does not mean the model’s weights changed.

It does not mean the model permanently learned the fact.

Ordinary inference still looks like:

fixed weights
+ current token context
-> next-token prediction

So context is more like temporary working material than permanent learning.

A finance analogy would be the documents open on my desk.

The model’s trained weights are like everything already built into the analyst’s knowledge and formulas.

The context is like the annual report, spreadsheet, and notes currently open for this one task.

Why attention seems to create a lot of repeated work

Earlier in AI Foundations we learned that attention creates query, key, and value vectors.

For one token representation, a layer computes something conceptually like:

hidden state
-> Q
-> K
-> V

When the model generates the next token, that new token’s query needs to compare against keys from earlier tokens.

Very roughly:

new query
compares with
keys from token 1
keys from token 2
keys from token 3
...

Then the attention weights are used to combine the corresponding values.

Without caching, the runtime could recompute the keys and values for all earlier tokens every time a new token is generated.

That would repeat a huge amount of work.

The key observation: old K and V do not need to change

During ordinary autoregressive decoding, the earlier tokens stay the same.

Suppose the context is:

The bond price fell because rates

and the model generates:

rose

The representations of the earlier tokens have already passed through the model for this sequence.

At each attention layer, the keys and values for those earlier positions have already been computed.

So instead of recalculating them at every decode step, the runtime can save them.

That saved data is the KV cache.

What KV cache actually means

KV stands for:

K = keys
V = values

The runtime stores the key and value tensors produced for previous tokens at the attention layers.

Then when a new token arrives, the model only needs to compute the new token’s fresh information and combine it with the cached history.

Conceptually:

step 1:
process prompt
store K/V for prompt tokens

step 2:
generate token A
store K/V for token A

step 3:
generate token B
reuse old K/V + add K/V for token B

step 4:
generate token C
reuse old K/V + add K/V for token C

The cache grows with the sequence.

My spreadsheet analogy

Suppose I have a financial model with 5,000 historical transactions.

I calculate several intermediate columns for every row:

normalized amount
category exposure
running balance
risk bucket

Now one new transaction arrives.

A terrible implementation would throw away all 5,000 intermediate results and recalculate them from the raw data.

A sensible implementation keeps the old intermediate results and calculates only what is needed for the new row.

KV cache feels similar.

The earlier token work is not all discarded after each generated token.

The runtime keeps reusable attention state.

Prefill and decode are different phases

This also explains two words that appear constantly in inference benchmarks:

prefill
decode

Prefill

During prefill, the model processes the input prompt.

If the prompt contains 4,000 tokens, the runtime needs to process those tokens through the transformer and construct the initial KV cache.

So prefill is roughly:

prompt
-> transformer computation
-> initial KV cache
-> first next-token logits

Decode

After that, generation proceeds one token at a time.

During decode, each new token can reuse the KV cache created from all previous positions.

So decode is roughly:

one new token
+ existing KV cache
-> one more token
+ enlarged KV cache

That is why benchmark results often report prompt-processing speed and token-generation speed separately.

They are different workloads.

Why the cache contains K and V but not every Q

This confused me initially.

If attention has Q, K, and V, why do we call it a KV cache instead of a QKV cache?

The reason is that the new token’s query is used to ask about the existing history.

For the next decoding step, we need:

new Q
against
old K values

and then use the matching old V values.

The earlier queries are not needed in the same reusable way for future tokens.

So the useful persistent state is primarily:

past K
past V

Hence:

KV cache

A tiny example

Imagine a toy sequence with three tokens:

The market fell

After processing them, one attention layer has cached:

K1, V1   for "The"
K2, V2   for "market"
K3, V3   for "fell"

Now the next generated token is being predicted.

The model computes a query for the newest position and compares it against:

K1
K2
K3

Then suppose the model generates:

today

The layer computes:

K4, V4

and appends them to the cache.

Now the next step can use:

K1, V1
K2, V2
K3, V3
K4, V4

The cache grows one position at a time.

Why long context consumes so much memory

The speed benefit has a cost.

We are storing key and value tensors for many layers and many tokens.

A simplified memory relationship is:

KV memory
roughly proportional to
number of tokens
x number of layers
x KV dimensions
x bytes per element

For a conventional dense transformer, a rough per-token expression is:

2 x layers x KV heads x head dimension x bytes per value

The 2 is for:

keys + values

Real architectures complicate this with grouped-query attention, multi-query attention, quantized cache formats, hybrid attention, sliding windows, and implementation details.

But the direction is still useful:

longer context -> larger KV cache
more concurrent sequences -> more total KV memory
higher-precision cache -> more bytes

This is why a model that fits easily in VRAM at short context can run into memory pressure when the context or concurrency becomes large.

Model weights and KV cache are different memory consumers

I used to think GPU memory was basically:

model size

But inference memory is more like several categories:

model weights
+ KV cache
+ activations / temporary workspace
+ runtime overhead

The model weights are mostly fixed after loading.

The KV cache is dynamic.

It grows and shrinks with active sequences.

That means two servers running the exact same model can have very different memory requirements depending on:

context length
number of concurrent requests
KV cache precision
serving strategy

Why concurrency makes the problem harder

Suppose one chat uses a KV cache of size X.

With ten active chats, the runtime may need cache state for ten independent token histories.

The exact allocation is implementation-dependent, but conceptually:

sequence A has its own history
sequence B has its own history
sequence C has its own history
...

The server cannot simply mix all of those cached states together because each request has a different context.

So long context and high concurrency compete for the same memory budget.

This is one reason serving systems care so much about KV-cache allocation efficiency.

PagedAttention: treat KV memory more like virtual memory

The vLLM PagedAttention work made this problem easier for me to picture.

Instead of requiring every sequence’s cache to occupy one perfectly sized contiguous region, the runtime can manage KV cache in blocks or pages.

Conceptually:

sequence A -> blocks 1, 4, 9
sequence B -> blocks 2, 3
sequence C -> blocks 5, 6, 7, 8

The physical layout does not have to match the logical token order exactly.

This is similar in spirit to operating-system virtual memory:

logical address space
!=
one giant contiguous physical allocation

The purpose is not to change what attention means.

It is to use memory more efficiently while serving many changing sequences.

Now I can state the difference clearly.

Context window

Answers:

How much token history may the model use?

It is fundamentally a model/runtime sequence-length limit.

KV cache

Answers:

How do we store reusable attention state from that history during inference?

It is an inference optimization and memory structure.

So:

context = what history is available
KV cache = saved K/V state for efficiently using that history

A larger context window usually creates the possibility of a larger KV cache, but they are not synonyms.

The cache does not mean the model remembers everything equally well

Another important distinction:

If a token is physically present in context and represented in the KV cache, that does not guarantee the model will use it perfectly.

The model still has to attend to the relevant information.

Long-context quality can depend on:

training
position handling
attention architecture
prompt structure
retrieval difficulty
runtime correctness

So this statement is too strong:

"It is in KV cache, therefore the model remembers it perfectly."

The cache preserves computational state.

It does not guarantee perfect reasoning or retrieval over that state.

Why clearing a chat can free memory

When an inference server no longer needs a sequence, its KV-cache allocation can be released or reused.

That is why ending a session, evicting an idle sequence, or reducing concurrency can free substantial serving memory even though the model weights stay loaded.

Conceptually:

model weights remain
sequence KV state disappears

This is also why server schedulers care about which requests remain active.

Prefix caching takes reuse one step further

Normal KV caching reuses state inside one sequence as it grows.

Some runtimes can also reuse cache blocks across requests that share an identical prefix.

For example, many requests may begin with the same long system prompt:

same 10,000-token system prefix
+ different user question

If the runtime can safely reuse the computed prefix state, it may avoid recomputing that identical portion for every request.

That is usually called prefix caching or prompt caching.

The important distinction is:

KV cache:
reuse earlier state within ongoing inference

prefix caching:
reuse matching earlier state across compatible requests/sessions

The implementation details vary, but the underlying motivation is the same: do not recompute expensive transformer work when the necessary state is already available.

A complete generation picture

At this point the inference path looks much clearer to me.

Start with a prompt:

text
-> tokenizer
-> token IDs

Then prefill:

token IDs
-> embeddings
-> transformer layers
-> build KV cache
-> logits

Then sample:

logits
-> temperature / top-k / top-p
-> choose token

Then decode:

new token
+ existing KV cache
-> next forward step
-> append new K/V
-> new logits
-> sample again

Repeat until the model stops or the generation limit is reached.

That connects the last several lessons into one continuous mechanism.

What I understand now

My old mental model was:

LLM reads the entire conversation again for every new word

That is incomplete.

A better model is:

prefill processes the existing context
KV cache stores reusable K/V state
new tokens are decoded incrementally
context and cache grow together

This explains several practical facts at once:

why long prompts have prefill cost
why later tokens can reuse earlier computation
why long context consumes more memory
why concurrent chats consume more KV capacity
why serving runtimes invest heavily in cache management

The next question follows naturally.

If model weights and KV cache consume so much memory, can we store some of those numbers with fewer bits?

That leads directly to quantization.

Next: AI Foundations #18 — Quantization: How Models Use Fewer Bits to Fit in Less Memory.

Sources and further reading

Continue reading