AI Fundamentals

What Is a KV Cache? How It Works and Why It Uses VRAM

Learn what an LLM KV cache stores, why it speeds up token generation, how to estimate its VRAM use, and how context length and quantization affect memory.

Approximately 15 min read


If you run large language models locally, you will eventually encounter the term KV cache.

It often appears when discussing:

The short explanation is:

A KV cache stores reusable attention information from tokens the model has already processed, so the model does not have to recompute that information every time it generates the next token.

That makes text generation much faster.

The trade-off is memory.

As the active context becomes longer, the KV cache generally becomes larger. For long-context inference or multiple simultaneous requests, KV-cache memory can consume several gigabytes of GPU memory even when the model weights themselves already fit.

KV cache in one example

Suppose you give a model this prompt:

The capital of France is

The model processes those tokens and generates:

Paris

Now it needs to generate another token.

Without a cache, the model would repeatedly recompute attention information for:

The
capital
of
France
is
Paris

Then it would repeat much of that work again for the next generated token.

With a KV cache, the model keeps the reusable attention state for the tokens it has already processed.

For the next generation step, it mainly needs to compute new attention state for the latest token and combine it with the cached state.

Conceptually:

Prompt tokens
    ↓
Compute K and V
    ↓
Store them in KV cache
    ↓
Generate next token
    ↓
Compute K and V for new token
    ↓
Append to cache
    ↓
Generate next token
    ↓
Repeat

This is why the cache grows as the sequence grows.

What do K and V mean?

Transformer attention works with three important representations:

Q = Query
K = Key
V = Value

A simplified attention operation can be written as:

Attention(Q, K, V)

The query represents what the current token is looking for.

The keys help determine which previous token representations are relevant.

The values contain information that can be retrieved after the attention scores determine what should receive attention.

For every transformer attention layer, the model projects token representations into Q, K, and V tensors.

During autoregressive generation, previously computed keys and values remain useful for future tokens.

That is why they are cached.

Why is it called a KV cache instead of a QKV cache?

This is an important detail.

When the model generates a new token, it needs the new token’s query to determine how strongly it should attend to previous tokens.

But the queries from previous generation steps do not need to be reused in the same way.

The previous tokens’ keys and values do.

So conceptually, at generation step t:

Current query:
q_t

Cached keys:
k_1, k_2, ... k_(t-1)

Cached values:
v_1, v_2, ... v_(t-1)

New key:
k_t

New value:
v_t

The model can calculate attention using:

q_t
against
[k_1 ... k_t]

and retrieve from

[v_1 ... v_t]

Then k_t and v_t become part of the cache for the next generation step.

Past queries do not need to be stored for this purpose.

Hence:

KV cache, not QKV cache.

What happens without a KV cache?

Autoregressive language models generate text one token at a time.

Imagine a model has already processed 10,000 tokens and is generating token 10,001.

Without caching, it would repeatedly calculate key and value projections for tokens it has already processed.

Then when generating token 10,002, much of that work would happen again.

The longer the sequence becomes, the more wasteful this gets.

A KV cache avoids repeatedly recomputing those past key and value tensors.

So the trade-off is:

Without KV cache:
less persistent cache memory
more repeated computation

With KV cache:
more memory
much less repeated computation

For normal autoregressive LLM inference, that is usually a very favorable trade.

Prefill and decode

LLM inference is often divided into two stages:

Prefill
Decode

Understanding them makes the KV cache much easier to understand.

Prefill

Suppose you submit a prompt containing 5,000 tokens.

During prefill, the model processes that prompt and constructs the KV state needed for those tokens.

Unlike token-by-token generation, many prompt tokens can be processed in parallel.

At the end of prefill, the model has a cache representing the relevant attention state for the prompt.

Decode

Now generation begins.

The model produces one new token.

Its new key and value tensors are appended to the cache.

Then another token is generated.

Again:

new token
→ compute new K and V
→ append to cache
→ generate next token

This continues until generation stops or the runtime reaches some context-management limit.

Does KV cache contain the text of your conversation?

Not directly.

The KV cache is not simply:

["Hello", "how", "are", "you"]

and it is not a copy of the prompt stored as plain text.

It contains numerical tensors derived from the tokens at each attention layer.

The runtime may separately retain token IDs or request information for other reasons, but those are conceptually different from the attention KV cache itself.

The KV cache is working inference state.

It is not:

When a request or inference session is discarded, its ordinary KV state can be discarded as well.

Why does KV cache use so much VRAM?

Because the model stores keys and values for many tokens across many transformer layers.

For a conventional decoder transformer, a useful approximate formula is:

KV cache bytes
≈
2
× number of layers
× number of tokens
× number of KV heads
× head dimension
× bytes per cache element
× number of active sequences

The leading 2 represents:

K
+
V

This formula is simplified, but it explains the major variables.

Example: a KV cache can easily reach 4 GiB

Consider a hypothetical transformer with:

Layers:          32
KV heads:         8
Head dimension: 128
Context:      32,768 tokens
KV dtype:      FP16
Active sequences: 1

FP16 requires two bytes per element.

The approximate cache size is:

2 × 32 × 32,768 × 8 × 128 × 2

which equals:

4,294,967,296 bytes

or approximately:

4 GiB

That is 4 GiB just for the KV cache of one sequence in this simplified example.

It does not include:

This is why a model whose weights occupy 18 GB can still run into trouble on a 24 GB GPU during a long-context request.

For a broader memory-planning guide, see How Much VRAM Do You Need for Local LLMs?.

Context length and KV-cache memory

In conventional full attention, KV-cache memory generally grows roughly linearly with the number of cached tokens.

Using our hypothetical example:

4K context   → roughly 0.5 GiB
8K context   → roughly 1 GiB
16K context  → roughly 2 GiB
32K context  → roughly 4 GiB

These numbers apply only to that example architecture.

The important relationship is:

2× tokens
≈
2× KV-cache storage

when the other variables remain constant.

This is one reason setting a huge context window can dramatically change memory requirements.

Maximum context is not the same as allocated KV-cache memory

There is an important runtime detail here.

If a model advertises:

128K context

that does not necessarily mean every runtime immediately allocates a full 128K cache every time you start the model.

Cache implementations differ.

For example, a dynamic cache can grow as tokens are added.

A static cache may preallocate storage for a predetermined maximum size.

Serving engines may allocate memory in blocks or pages and manage those blocks across requests.

Therefore:

model supports 128K context

does not automatically tell you exactly how much VRAM will be allocated at startup.

You need to consider both the model architecture and the runtime’s cache implementation.

Why GQA can dramatically reduce KV-cache size

Not all transformers have the same number of query and KV heads.

With traditional Multi-Head Attention (MHA), the model may use the same number of query, key, and value heads.

For example:

32 query heads
32 key heads
32 value heads

Many modern LLMs instead use Grouped-Query Attention (GQA).

A GQA model might have:

32 query heads
8 KV heads

Multiple query heads share the same key/value heads.

Because the cache stores K and V rather than Q, reducing the number of KV heads can substantially reduce cache memory.

In this example:

32 KV heads
→ baseline

8 KV heads
→ approximately one-quarter the K/V tensor storage

assuming the other relevant dimensions are unchanged.

What is MQA?

Multi-Query Attention (MQA) takes this idea further.

Multiple query heads can share a single set of key/value heads.

Conceptually:

many query heads
1 KV head

This can make KV-cache requirements much smaller than conventional MHA.

That is why two models with the same parameter count and the same context length can have very different KV-cache sizes.

You cannot accurately estimate KV memory from parameter count alone.

Sliding-window attention changes the picture

Some models do not keep unrestricted full attention over every previous token in every layer.

With sliding-window attention, a layer may retain attention state only for a limited recent window.

Once that layer reaches its configured window, its cache no longer has to grow indefinitely with the full sequence.

Other architectures use chunked, hybrid, recurrent, state-space, or specialized attention mechanisms.

So the simple linear formula is most useful as a conceptual estimate for conventional cached self-attention.

Always inspect the architecture when precise memory planning matters.

KV cache vs model weights

These are separate memory consumers.

Suppose you run a quantized model.

Your VRAM might conceptually look like:

Model weights      16 GB
KV cache            4 GB
Runtime/buffers      2 GB
Other/headroom       2 GB
-------------------------
Total               24 GB

Reducing model weight size does not automatically reduce KV-cache size.

For example:

FP16 model
→ convert to 4-bit weights

can greatly reduce weight memory.

But if the KV cache remains FP16, its storage is still based on its own precision and dimensions.

This distinction is important when comparing quantization formats such as GGUF and AWQ.

Weight quantization and KV-cache quantization are different

Consider:

Model weights: 4-bit
KV cache: FP16

The model has quantized weights, but its cache is not 4-bit.

These are two separate decisions.

Modern inference engines may support lower-precision KV caches, including FP8 or other quantized representations.

That can substantially reduce cache memory.

For the hypothetical 4 GiB FP16 cache above, an idealized one-byte-per-element representation would reduce the raw tensor storage to roughly:

2 GiB

before considering implementation-specific overhead.

Does quantizing the KV cache affect quality?

Potentially.

Unlike lossless storage compression, changing numerical precision changes how values are represented.

A lower-precision KV cache can introduce quantization error.

Whether that matters noticeably depends on:

So the decision should not be reduced to:

Lower precision is always better.

The benefit is lower memory usage.

The trade-off can include numerical differences and, depending on the implementation and hardware, different performance characteristics.

KV-cache offloading

Another strategy is to move some KV-cache storage away from the GPU.

For example:

GPU
↕
CPU RAM

Cache offloading can reduce GPU-memory pressure and allow workloads that would otherwise exceed VRAM.

The trade-off is data movement.

GPU memory is usually much faster for GPU inference than transferring data through the CPU-memory path.

So offloading can exchange:

lower VRAM usage

for:

potentially higher latency

Hugging Face Transformers and current inference runtimes provide cache-offloading mechanisms for memory-constrained scenarios.

llama.cpp KV-cache controls

llama.cpp exposes explicit controls for KV-cache behavior.

Current command-line options include separate cache types for keys and values, such as:

--cache-type-k
--cache-type-v

and a KV-offload control.

This matters because a GGUF model’s weight quantization and the runtime’s KV-cache representation are independent choices.

You might therefore use:

Q4_K_M model weights

while choosing a different type for the runtime KV cache.

That distinction is easy to miss.

Why concurrency makes KV cache expensive

So far we have mainly discussed one sequence.

Now imagine a server handling multiple users simultaneously.

Each request can have its own active context.

Conceptually:

Request A → KV state
Request B → KV state
Request C → KV state
Request D → KV state

The model weights can be shared by all requests.

Their active KV state cannot simply be treated as one identical cache.

This is why KV-cache capacity is a major issue for LLM serving.

More available cache memory can allow a serving engine to keep more tokens and more concurrent sequences active.

So a GPU configuration that works beautifully for one person’s local chat may become memory constrained when used as a multi-user API server.

Why vLLM cares so much about KV cache

High-throughput serving engines such as vLLM actively manage KV-cache memory.

Rather than treating cache allocation as an afterthought, the inference engine organizes cache storage so it can serve multiple sequences efficiently.

Current vLLM configurations expose controls for concepts such as:

This reflects how important cache management becomes when maximizing throughput on expensive GPU memory.

What is prefix caching?

Prefix caching is related to ordinary KV caching, but it solves a different problem.

Normal KV caching avoids recomputing earlier tokens within a generation sequence.

Prefix caching can reuse previously computed KV state when multiple requests share the same prefix.

Imagine many requests begin with the same long system prompt:

You are an assistant for Acme Corporation.
Follow these 5,000 tokens of policies...

Without prefix reuse, each request may need to process that same prefix again.

With prefix caching, an inference engine can identify reusable prefix state and avoid repeating some of that prompt-processing work.

Conceptually:

Ordinary KV cache:
reuse past tokens inside one active sequence

Prefix cache:
reuse a previously computed common prefix across compatible requests

This can be particularly useful for API serving with large shared system prompts.

Does prefix caching mean users share conversations?

No.

Reusing computed state for identical compatible prefixes does not mean arbitrary private conversation histories become one shared conversation.

Serving systems need explicit cache-management and isolation logic.

The important conceptual point is simply that identical prefixes can represent duplicated computation that an optimized engine may avoid.

Why long-context models make KV cache important

Modern models can advertise very large context windows.

That allows use cases such as:

But model support for a context length does not mean the context is free.

Longer context can increase:

That is why context length is both a model capability and a hardware-planning variable.

Why a model can load successfully and later run out of memory

This is a very common local-LLM scenario.

You start the model:

Model loaded successfully.

VRAM usage appears acceptable.

Then you paste a large document.

Later:

CUDA out of memory

Nothing is contradictory here.

At startup, much of the fixed memory may come from model weights and runtime allocations.

As the request grows, additional KV-cache state is created.

The workload eventually crosses the available-memory limit.

So:

model fits

does not necessarily mean:

maximum context fits

Ways to reduce KV-cache memory

Depending on your runtime and model, the main options include:

Use a shorter context

This is usually the simplest solution.

If you do not need 128K tokens, do not allocate or retain 128K worth of active context unnecessarily.

Reduce concurrency

For server workloads, fewer simultaneously active sequences mean less aggregate cache state.

Use a lower-precision cache

If supported, FP8 or another quantized cache representation can reduce memory use.

Use KV-cache offloading

Moving cache data to CPU memory can free GPU VRAM at the cost of additional data movement.

Choose an architecture with fewer KV heads

GQA and MQA can significantly reduce KV-cache storage compared with otherwise similar full MHA architectures.

This is a model-selection decision rather than a runtime switch.

Use sliding-window or specialized attention when supported

Some architectures limit the amount of state retained by particular layers.

Again, this depends on the model architecture.

Does reducing KV cache make generation faster?

Not necessarily.

Memory capacity and inference speed are related, but they are not the same metric.

For example:

FP16 KV cache
→ larger memory footprint

FP8 KV cache
→ smaller memory footprint

The lower-memory representation may let you retain more tokens or support more concurrent requests.

But actual latency depends on whether the hardware and inference kernels efficiently support that representation.

Similarly:

GPU KV cache

may use more scarce GPU memory but avoid transfers, while:

CPU-offloaded KV cache

saves VRAM but can introduce additional movement and latency.

Always benchmark the configuration that matches your workload.

Is KV cache used during training?

KV caching is primarily an inference optimization for autoregressive generation.

Training works differently because the model must compute gradients and process attention states needed for backpropagation.

Reusing a normal inference KV cache during training would interfere with the assumptions of the training computation.

So when you see KV-cache controls in tools such as Transformers, llama.cpp, or vLLM, think primarily about inference.

A practical troubleshooting checklist

If your LLM runs out of VRAM during inference, check these variables before assuming the model itself is too large:

1. Model weight size
2. Weight quantization
3. Configured context length
4. Actual tokens currently in context
5. KV-cache precision
6. Number of KV heads
7. Number of simultaneous sequences
8. Runtime memory settings
9. GPU cache offloading options
10. Other processes consuming VRAM

This separates fixed model memory from context-dependent memory.

KV cache and VRAM: the most important mental model

Think of model weights as the relatively fixed part:

Model loaded
→ weight memory mostly established

Think of KV cache as active working memory:

conversation gets longer
→ more attention state
→ larger cache

For multiple users:

more active sequences
→ more aggregate cache state

That simple distinction explains a large portion of seemingly mysterious LLM VRAM behavior.

Bottom line

A KV cache stores the key and value tensors from previously processed tokens at the model’s attention layers.

Its purpose is to avoid repeatedly recomputing those K/V states during autoregressive generation.

That produces a fundamental trade-off:

More KV caching
→ less repeated computation
→ more memory use

The amount of memory depends primarily on:

number of layers
× cached tokens
× KV heads
× head dimension
× cache precision
× active sequences

Long context and high concurrency can therefore consume large amounts of VRAM even when model weights fit comfortably.

Weight quantization and KV-cache quantization are separate optimizations.

And when diagnosing LLM memory problems, remember:

The question is not only whether the model fits in VRAM. The context has to fit too.

Sources and further reading

Continue reading