Troubleshooting

How to Reduce LLM VRAM Usage Without Breaking Performance

Learn how to reduce local LLM VRAM usage with quantization, shorter context, KV-cache compression, CPU offloading, smaller models, and runtime tuning.

Approximately 18 min read

You try to load a local LLM and get:

CUDA out of memory

or the model loads successfully but fails as soon as you submit a long prompt.

The obvious conclusion is:

I need a bigger GPU.

Sometimes that is true.

But there are several ways to reduce LLM VRAM usage before buying new hardware.

The most important memory consumers are usually:

model weights
+
KV cache
+
runtime buffers
+
temporary allocations

So the solution depends on which part is consuming the memory.

A useful troubleshooting order is:

1. Use a smaller weight quantization
2. Reduce context length
3. Reduce KV-cache precision
4. Move KV cache to CPU RAM
5. Offload fewer model layers to GPU
6. Reduce concurrency
7. Reduce batch-related memory
8. Choose a smaller model

Each method has a different trade-off.

Some preserve model quality but reduce speed.

Others preserve speed but reduce context.

Others change the model itself.

This guide explains when to use each one.

First understand where VRAM goes

A simplified inference memory equation is:

Total VRAM
≈
model weights
+
KV cache
+
runtime buffers
+
temporary working memory

If you are running an image model, training job, browser GPU workload, or another LLM at the same time, those also compete for VRAM.

So before changing anything, inspect actual GPU usage.

On NVIDIA systems:

nvidia-smi

For continuous monitoring:

watch -n 1 nvidia-smi

Look at:

total VRAM
used VRAM
free VRAM
other GPU processes

This tells you whether the model itself is the problem or whether another application is occupying several gigabytes.

Model weights are usually the largest fixed cost

Suppose you load a:

32B model

Theoretical weight storage is approximately:

FP16
32B × 2 bytes
≈ 64 GB

8-bit
32B × 1 byte
≈ 32 GB

4-bit
32B × 0.5 byte
≈ 16 GB

Real quantized models require additional metadata, so actual files do not exactly match these numbers.

Still, the calculation shows why quantization is the first major VRAM optimization.

If your weights alone are too large, reducing context will not solve the entire problem.

Fix 1: use a smaller quantization

Suppose you currently use:

Q8_0

and it consumes too much memory.

You might try:

Q6_K
Q5_K_M
Q4_K_M

depending on what is available.

Conceptually:

Q8
→ larger

Q6
→ smaller

Q5
→ smaller

Q4
→ smaller again

Lower-bit quantization stores model weights more compactly.

That can free several gigabytes of VRAM.

For a detailed comparison, read GGUF Quantization Levels Explained: Q4, Q5, Q6 and Q8.

Quantization does not reduce parameter count

A:

32B Q4 model

is still approximately a:

32-billion-parameter model

The model contains the same general parameter count.

Those parameters are simply stored using a lower-precision representation.

So quantization reduces:

memory per weight

not:

number of weights

For the underlying concepts, read What Is LLM Quantization? FP16, INT8 and INT4 Explained.

Example: Q8 vs Q4

Imagine the same model is available as:

Q8_0:
30 GB

Q4_K_M:
17 GB

Your GPU has:

24 GB VRAM

The Q8 model cannot fit fully in VRAM.

The Q4 model may fit while leaving room for:

KV cache
runtime buffers
context

That difference can completely change how the model runs.

Smaller quantization can sometimes improve speed

This is possible because lower-bit weights reduce memory traffic and may allow the entire model to fit on the GPU.

For example:

Q6
→ partially CPU offloaded

Q4
→ fully GPU resident

The Q4 configuration may be much faster.

But this is not guaranteed.

Actual performance depends on:

runtime
GPU architecture
CPU
memory bandwidth
quantization kernels
model architecture

So do not assume:

smaller file
=
always faster

Measure it.

Fix 2: reduce context length

If the model loads successfully but later runs out of memory, the problem may be the KV cache.

Context length is one of the largest variable memory costs during LLM inference.

Suppose you configured:

128K context

but you normally use only:

5K–10K tokens

You may be reserving or allowing far more context capacity than the workload requires.

In llama.cpp, context is controlled with:

-c
--ctx-size

For example:

llama-cli \
  -m model.gguf \
  -c 8192

Instead of:

llama-cli \
  -m model.gguf \
  -c 65536

if you do not actually need 65K tokens.

Why shorter context saves memory

Autoregressive inference commonly stores keys and values for previous tokens in the:

KV cache

For conventional full attention, the raw amount of cached K/V state grows approximately with the number of cached tokens.

Conceptually:

4K context
→ smaller cache

8K
→ larger

16K
→ larger again

32K
→ larger again

The exact memory depends heavily on architecture.

Models using:

GQA
MQA
sliding-window attention
chunked attention

can behave differently.

For the full explanation, read What Is a KV Cache? How It Works and Why It Uses VRAM.

Maximum context is not a target

If a model says:

128K context supported

that does not mean you should always configure 128K.

Context capacity is a maximum capability.

A better rule is:

use enough context for the workload
+
leave memory headroom

If your workload fits in:

8K

then 8K may be a much more practical local configuration than 128K.

Read What Does Context Length Mean in an LLM? for more detail.

Fix 3: reduce KV-cache precision

Model weights and KV cache can use different numerical representations.

You could have:

weights:
Q4_K_M

KV cache:
FP16

This means you have already compressed the model weights but may still be spending significant VRAM on the cache.

Some runtimes allow lower-precision KV-cache storage.

Current llama.cpp exposes:

-ctk
--cache-type-k

and:

-ctv
--cache-type-v

for key and value cache types.

The default is currently:

f16

and supported options include lower-precision formats such as:

q8_0
q5_0
q5_1
q4_0
q4_1
iq4_nl

depending on your build and model.

Example: llama.cpp KV-cache quantization

A test configuration might be:

llama-cli \
  -m model.gguf \
  -c 32768 \
  -ngl auto \
  -ctk q8_0 \
  -ctv q8_0

This changes KV-cache representation independently from model-weight quantization.

You can test more aggressive cache types if supported.

But do not change directly from FP16 to the smallest possible cache representation and assume the result is free.

Lower-precision KV cache has trade-offs

KV-cache quantization can reduce memory.

But possible costs include:

additional quantize/dequantize work
latency changes
numerical approximation
model-specific quality effects

For short contexts where VRAM is already sufficient, cache quantization may provide little benefit and can even hurt latency.

Use it primarily when KV-cache memory is genuinely the bottleneck.

Weight quantization and KV quantization are separate

This distinction is worth repeating.

You can have:

Q4 weights
+
FP16 KV

or:

Q4 weights
+
Q8 KV

or another supported combination.

Changing the GGUF quantization does not automatically change the KV-cache format.

This is why a heavily quantized model can still consume substantial VRAM at long context.

Fix 4: move KV cache from GPU to CPU RAM

If VRAM is extremely constrained but system RAM is available, the KV cache can sometimes be moved to CPU memory.

This saves GPU VRAM.

The trade-off is additional communication between CPU memory and GPU processing.

In current llama.cpp, GPU KV offloading is enabled by default.

To keep KV-cache buffers in CPU RAM, you can use:

--no-kv-offload

or:

-nkvo

For example:

llama-cli \
  -m model.gguf \
  -c 32768 \
  -ngl all \
  --no-kv-offload

This can make a configuration fit when GPU memory would otherwise be exhausted.

CPU KV cache can hurt performance

Moving KV state away from the GPU means the runtime must rely more heavily on CPU memory and data transfer.

The model may still work.

But prompt processing and generation can become slower.

So:

KV cache on GPU
→ more VRAM
→ generally faster

KV cache in system RAM
→ less VRAM
→ potentially slower

This is a classic memory-versus-latency trade-off.

Hugging Face cache offloading

Transformers also supports cache offloading.

The general principle is:

most KV cache
→ CPU

currently needed layer
→ GPU

The runtime transfers cache state as layers are processed.

This saves VRAM while preserving the ability to generate long sequences.

Again, the expected trade-off is lower GPU memory usage at the cost of additional transfer overhead.

Fix 5: offload fewer model layers to the GPU

If the model weights themselves do not fit, llama.cpp can split execution between:

GPU VRAM

and:

system RAM / CPU

Current llama.cpp controls GPU layer placement with:

-ngl
--n-gpu-layers

You may use:

auto
all
or an exact layer count

For example:

llama-cli \
  -m model.gguf \
  -ngl auto

lets the runtime choose placement.

If you deliberately need to reduce VRAM further, you can test a lower exact value.

For example:

llama-cli \
  -m model.gguf \
  -ngl 20

The correct number depends on model architecture.

Do not copy someone else’s layer count blindly.

Partial GPU offloading

Conceptually:

GPU:
layers 1–20

CPU / RAM:
remaining layers

This allows a model larger than available VRAM to run.

The trade-off is performance.

More work on the CPU generally means lower generation speed than a fully GPU-resident model.

Full GPU vs hybrid inference

A simplified comparison:

Full GPU
→ more VRAM required
→ usually faster

Hybrid CPU/GPU
→ less VRAM required
→ usually slower

CPU-only
→ minimal dedicated VRAM
→ often much slower

The best point depends on:

GPU
CPU
RAM bandwidth
model size
quantization

Fix 6: use a smaller model

Sometimes you have already optimized everything else.

If your hardware cannot comfortably run:

70B

the most effective fix may be:

32B
14B
8B

rather than trying increasingly extreme memory tricks.

Parameter count has a huge effect on model-weight memory.

For example, idealized 4-bit dense weights require approximately:

8B
≈ 4 GB

14B
≈ 7 GB

32B
≈ 16 GB

70B
≈ 35 GB

before quantization metadata and runtime allocations.

A smaller model can free more VRAM than almost any individual runtime setting.

Bigger is not always better for your workload

A larger model can provide better capability.

But if it forces:

heavy CPU offloading
tiny context
severe quantization
constant OOM risk
very low tokens/s

a smaller model running efficiently may provide a better user experience.

The real objective is:

quality
+
speed
+
memory

not simply the largest parameter count possible.

Fix 7: reduce parallel sequences

Single-user chat and multi-user serving have very different memory requirements.

Each active sequence requires inference state.

If you serve:

1 user

you need much less cache state than:

16 concurrent users

In llama.cpp, parallel decoding is controlled with:

-np
--parallel

For a personal workstation:

-np 1

may be enough.

If you do not need concurrency, do not allocate for it.

vLLM and concurrency

Serving engines such as vLLM dedicate substantial GPU memory to model execution and KV-cache capacity.

Current vLLM exposes controls such as:

gpu_memory_utilization
kv_cache_memory_bytes
cache_dtype
kv_offloading_size

These let administrators balance:

model memory
KV-cache capacity
concurrency
VRAM limits

A production server optimized for maximum concurrent requests should not be configured the same way as a single-user desktop.

Do not blindly maximize gpu_memory_utilization

A serving engine may allow you to reserve a high fraction of GPU memory.

That does not mean every machine should run at the maximum possible value.

Leave room for:

driver overhead
other GPU workloads
monitoring tools
multiple model instances
unexpected allocations

If the server is repeatedly near OOM, memory headroom is more useful than squeezing out one additional cache block.

Explicit KV-cache limits in vLLM

Current vLLM can automatically infer KV-cache capacity from its GPU-memory allocation.

It also exposes:

kv_cache_memory_bytes

for more direct control.

This can be useful when you want to reserve a known memory budget rather than allow cache allocation to consume whatever remains.

The right value depends on:

model size
maximum context
concurrency
GPU capacity

vLLM KV-cache dtype

vLLM also supports configuring the KV-cache data type.

Using a lower-precision cache can reduce VRAM requirements.

Again, this is independent from weight quantization.

Conceptually:

AWQ / GPTQ / FP8 weights

and:

KV-cache dtype

are separate settings.

vLLM KV-cache offloading

Current vLLM also supports CPU KV-cache offloading.

Its cache configuration exposes:

kv_offloading_size

which enables a CPU offloading buffer when configured.

This is useful when:

GPU KV capacity is insufficient

but:

system RAM is available

As with other offloading strategies, data movement can affect latency.

Prompt processing can require temporary working memory.

In llama.cpp, relevant controls include:

-b
--batch-size

and:

-ub
--ubatch-size

If you encounter OOM during large prompt processing rather than during steady generation, lowering these values can sometimes help.

For example:

llama-cli \
  -m model.gguf \
  -b 512 \
  -ub 256

instead of much larger settings.

Do not change these unless you have evidence that batch working memory is the problem.

Smaller batches can reduce speed

Larger batches can improve prompt-processing throughput on suitable hardware.

So reducing batch size may trade:

lower peak memory

for:

slower prefill

This is another reason to diagnose the failure before changing settings.

Model loads but prompt causes OOM

This scenario strongly suggests:

weights fit

but:

context / cache / temporary memory does not

First try:

shorter context

For example:

-c 8192

instead of:

-c 32768

If that fixes the problem, model-weight quantization may not be the main issue.

OOM immediately during model loading

If the model fails before inference even begins, investigate:

model weights
GPU-layer placement
other GPU applications
runtime buffers

The strongest interventions are usually:

smaller quantization
fewer GPU layers
smaller model

Context reduction may not matter much if the weights themselves cannot be loaded.

OOM after a long conversation

If the system works initially but fails after many conversation turns, investigate:

context growth
KV-cache growth
conversation history

Start a fresh chat.

If the new conversation works, context was probably a major factor.

OOM only with multiple users

If one request works but many simultaneous requests fail, investigate:

concurrency
parallel sequence count
KV-cache capacity
batching

This is a serving-capacity problem.

Reducing model quantization may help, but the more direct solution may be reducing concurrency or cache requirements.

Close other GPU programs

Do not overlook the simplest fix.

Your GPU may already contain allocations from:

browser
desktop compositor
image generator
video generator
another LLM
PyTorch notebook
game
CUDA process

Check:

nvidia-smi

before launching the model.

Freeing:

2–4 GB

from unrelated applications may be enough to avoid OOM.

Do not use all VRAM intentionally

If your GPU has:

24 GB

do not design around:

23.99 GB

of normal usage.

You need headroom.

Model memory can vary with:

prompt length
runtime version
temporary buffers
concurrency
driver behavior

A stable configuration is better than one that survives only one carefully controlled benchmark.

How much headroom should you leave?

There is no universal percentage.

The correct amount depends on the runtime and workload.

Instead of a fixed rule, test the largest realistic workload:

longest prompt
largest expected context
maximum concurrency
normal desktop applications

Then verify that the system remains stable.

Weight memory vs KV-cache memory

A practical diagnosis is:

If VRAM is already almost full immediately after model load:
→ weight problem

If VRAM grows significantly with context:
→ KV-cache problem

If OOM appears only during prompt processing:
→ temporary/batch memory may matter

If OOM appears only under many requests:
→ concurrency/cache problem

This is much more useful than changing every setting at once.

A good llama.cpp low-VRAM baseline

Suppose a model nearly fits.

Start with something conservative:

llama-cli \
  -m model.gguf \
  -c 8192 \
  -ngl auto \
  --perf

Then monitor:

nvidia-smi

If memory remains too high, try the optimizations one at a time.

Step 1: reduce quantization

For example:

Q6_K
↓
Q5_K_M
↓
Q4_K_M

Reload and measure.

Step 2: reduce context

For example:

-c 32768

to:

-c 8192

Reload and measure.

Step 3: reduce KV precision

Test:

-ctk q8_0 -ctv q8_0

instead of the default FP16 cache.

Measure:

VRAM
quality
prompt speed
generation speed

Step 4: move KV cache to CPU

If VRAM is still insufficient:

--no-kv-offload

This can save GPU memory but may reduce performance.

Step 5: reduce GPU layers

If the weights still do not fit:

-ngl 20

or another model-appropriate count.

Allow more of the model to remain in system RAM.

Change one variable at a time

Do not simultaneously change:

quantization
context
GPU layers
KV precision
batch size

and then conclude:

it is faster now

You will not know which change mattered.

A controlled optimization workflow is:

baseline
↓
change one setting
↓
measure
↓
keep or revert
↓
change next setting

Measure both VRAM and speed

A memory optimization is not automatically a good optimization.

Suppose:

Configuration A
VRAM: 22 GB
Speed: 40 tokens/s

Configuration B
VRAM: 15 GB
Speed: 8 tokens/s

Configuration B uses less memory.

But if you already have 24 GB available, the performance loss may not be worthwhile.

The target is usually:

fit comfortably

not:

minimize VRAM at any cost

Measure quality too

Aggressive weight quantization and KV-cache quantization can affect numerical behavior.

If your workload is:

coding
reasoning
structured extraction
long-context retrieval

test real examples.

Do not optimize only against:

nvidia-smi

The model still needs to solve the task.

Optimization priority by cost

A useful mental ordering is:

Close other GPU apps
→ essentially free

Reduce unused context
→ usually low quality cost

Choose a sensible model quantization
→ moderate trade-off

Reduce KV-cache precision
→ memory/latency/quality trade-off

Move KV cache to CPU
→ speed trade-off

Offload model layers to CPU
→ potentially large speed trade-off

Choose a smaller model
→ changes model capability

The best solution is usually the least disruptive change that makes the workload fit.

Example: 24 GB GPU with a 32B model

Suppose:

GPU:
24 GB

Model:
32B Q5_K_M

Context:
32K

and you get OOM.

A sensible sequence is:

1. Reduce context to 8K
2. Test again
3. If still too large, try Q4_K_M
4. Test again
5. Reduce KV-cache precision if needed
6. Move KV cache to CPU only if necessary

Do not immediately switch to CPU-only inference.

Example: model loads but 64K context fails

Suppose:

Model loaded:
20 GB VRAM

64K prompt:
OOM

This strongly suggests that the remaining:

4 GB

was insufficient for cache and working allocations.

Possible fixes:

32K context
16K context
quantized KV cache
CPU KV cache
smaller weight quantization

The model itself may be fine.

Example: 70B on one 24 GB GPU

A conventional dense 70B model at an idealized four bits already requires roughly:

35 GB

for weights before overhead.

Therefore it cannot normally fit entirely in 24 GB VRAM.

Your choices include:

CPU/GPU hybrid inference
multiple GPUs
more aggressive quantization
smaller model
larger-memory GPU

Runtime tuning cannot make 35 GB of weights magically become 24 GB without changing representation or placement.

Multiple GPUs

Multiple GPUs can reduce the memory pressure on any single GPU by distributing model tensors and, depending on runtime, cache state.

llama.cpp supports multi-GPU split modes and tensor distribution controls.

However:

two GPUs

do not automatically behave like:

one perfectly unified pool

Interconnect bandwidth, split strategy, backend, and architecture matter.

Multi-GPU is a deployment strategy, not a free memory expansion.

Unified memory systems

On Apple Silicon, CPU and GPU share unified memory.

The problem becomes less about:

dedicated VRAM

and more about:

total available unified memory

Quantization and context still matter because model weights, KV cache, applications, and the operating system all compete for the same pool.

You still need headroom.

VRAM optimization checklist

When local inference runs out of memory, check:

Is another application using the GPU?

How large are the model weights?

What quantization am I using?

What context size is configured?

How large is the active conversation?

What KV-cache precision is being used?

Is KV cache stored on GPU?

How many layers are on the GPU?

How many sequences are active?

Are batch settings unusually large?

Answer those questions before buying hardware.

What should you change first?

For most local users, a sensible sequence is:

1. Close unnecessary GPU applications

2. Use a realistic context length

3. Choose a practical model quantization

4. Reduce KV-cache precision if long context is the issue

5. Move KV cache to CPU if VRAM is still insufficient

6. Reduce GPU layer count

7. Choose a smaller model if performance becomes unacceptable

This preserves performance as long as possible.

Bottom line

Reducing LLM VRAM usage is not one trick.

VRAM is consumed by several components:

model weights
KV cache
runtime buffers
temporary memory
parallel sequences

Different optimizations target different components.

Use:

lower-bit model quantization

to reduce weight memory.

Use:

shorter context

to reduce cache requirements.

Use:

lower-precision KV cache

to compress long-context state.

Use:

CPU KV-cache placement

when VRAM is more important than latency.

Use:

partial GPU offloading

when the model itself does not fit.

And use:

a smaller model

when the compromises required to run a larger model make the system impractical.

The objective is not to achieve the lowest possible VRAM number.

It is to find a configuration where:

the model fits
+
the context fits
+
performance remains acceptable
+
quality remains acceptable
+
there is enough memory headroom for real workloads

That is the practical way to optimize local LLM VRAM usage.

Sources and further reading

Continue reading