How to Reduce LLM VRAM Usage Without Breaking Performance
Learn how to reduce local LLM VRAM usage with quantization, shorter context, KV-cache compression, CPU offloading, smaller models, and runtime tuning.
Approximately 18 min read
You try to load a local LLM and get:
CUDA out of memory
or the model loads successfully but fails as soon as you submit a long prompt.
The obvious conclusion is:
I need a bigger GPU.
Sometimes that is true.
But there are several ways to reduce LLM VRAM usage before buying new hardware.
The most important memory consumers are usually:
model weights
+
KV cache
+
runtime buffers
+
temporary allocations
So the solution depends on which part is consuming the memory.
A useful troubleshooting order is:
1. Use a smaller weight quantization
2. Reduce context length
3. Reduce KV-cache precision
4. Move KV cache to CPU RAM
5. Offload fewer model layers to GPU
6. Reduce concurrency
7. Reduce batch-related memory
8. Choose a smaller model
Each method has a different trade-off.
Some preserve model quality but reduce speed.
Others preserve speed but reduce context.
Others change the model itself.
This guide explains when to use each one.
First understand where VRAM goes
A simplified inference memory equation is:
Total VRAM
≈
model weights
+
KV cache
+
runtime buffers
+
temporary working memory
If you are running an image model, training job, browser GPU workload, or another LLM at the same time, those also compete for VRAM.
So before changing anything, inspect actual GPU usage.
On NVIDIA systems:
nvidia-smi
For continuous monitoring:
watch -n 1 nvidia-smi
Look at:
total VRAM
used VRAM
free VRAM
other GPU processes
This tells you whether the model itself is the problem or whether another application is occupying several gigabytes.
Model weights are usually the largest fixed cost
Suppose you load a:
32B model
Theoretical weight storage is approximately:
FP16
32B × 2 bytes
≈ 64 GB
8-bit
32B × 1 byte
≈ 32 GB
4-bit
32B × 0.5 byte
≈ 16 GB
Real quantized models require additional metadata, so actual files do not exactly match these numbers.
Still, the calculation shows why quantization is the first major VRAM optimization.
If your weights alone are too large, reducing context will not solve the entire problem.
Fix 1: use a smaller quantization
Suppose you currently use:
Q8_0
and it consumes too much memory.
You might try:
Q6_K
Q5_K_M
Q4_K_M
depending on what is available.
Conceptually:
Q8
→ larger
Q6
→ smaller
Q5
→ smaller
Q4
→ smaller again
Lower-bit quantization stores model weights more compactly.
That can free several gigabytes of VRAM.
For a detailed comparison, read GGUF Quantization Levels Explained: Q4, Q5, Q6 and Q8.
Quantization does not reduce parameter count
A:
32B Q4 model
is still approximately a:
32-billion-parameter model
The model contains the same general parameter count.
Those parameters are simply stored using a lower-precision representation.
So quantization reduces:
memory per weight
not:
number of weights
For the underlying concepts, read What Is LLM Quantization? FP16, INT8 and INT4 Explained.
Example: Q8 vs Q4
Imagine the same model is available as:
Q8_0:
30 GB
Q4_K_M:
17 GB
Your GPU has:
24 GB VRAM
The Q8 model cannot fit fully in VRAM.
The Q4 model may fit while leaving room for:
KV cache
runtime buffers
context
That difference can completely change how the model runs.
Smaller quantization can sometimes improve speed
This is possible because lower-bit weights reduce memory traffic and may allow the entire model to fit on the GPU.
For example:
Q6
→ partially CPU offloaded
Q4
→ fully GPU resident
The Q4 configuration may be much faster.
But this is not guaranteed.
Actual performance depends on:
runtime
GPU architecture
CPU
memory bandwidth
quantization kernels
model architecture
So do not assume:
smaller file
=
always faster
Measure it.
Fix 2: reduce context length
If the model loads successfully but later runs out of memory, the problem may be the KV cache.
Context length is one of the largest variable memory costs during LLM inference.
Suppose you configured:
128K context
but you normally use only:
5K–10K tokens
You may be reserving or allowing far more context capacity than the workload requires.
In llama.cpp, context is controlled with:
-c
--ctx-size
For example:
llama-cli \
-m model.gguf \
-c 8192
Instead of:
llama-cli \
-m model.gguf \
-c 65536
if you do not actually need 65K tokens.
Why shorter context saves memory
Autoregressive inference commonly stores keys and values for previous tokens in the:
KV cache
For conventional full attention, the raw amount of cached K/V state grows approximately with the number of cached tokens.
Conceptually:
4K context
→ smaller cache
8K
→ larger
16K
→ larger again
32K
→ larger again
The exact memory depends heavily on architecture.
Models using:
GQA
MQA
sliding-window attention
chunked attention
can behave differently.
For the full explanation, read What Is a KV Cache? How It Works and Why It Uses VRAM.
Maximum context is not a target
If a model says:
128K context supported
that does not mean you should always configure 128K.
Context capacity is a maximum capability.
A better rule is:
use enough context for the workload
+
leave memory headroom
If your workload fits in:
8K
then 8K may be a much more practical local configuration than 128K.
Read What Does Context Length Mean in an LLM? for more detail.
Fix 3: reduce KV-cache precision
Model weights and KV cache can use different numerical representations.
You could have:
weights:
Q4_K_M
KV cache:
FP16
This means you have already compressed the model weights but may still be spending significant VRAM on the cache.
Some runtimes allow lower-precision KV-cache storage.
Current llama.cpp exposes:
-ctk
--cache-type-k
and:
-ctv
--cache-type-v
for key and value cache types.
The default is currently:
f16
and supported options include lower-precision formats such as:
q8_0
q5_0
q5_1
q4_0
q4_1
iq4_nl
depending on your build and model.
Example: llama.cpp KV-cache quantization
A test configuration might be:
llama-cli \
-m model.gguf \
-c 32768 \
-ngl auto \
-ctk q8_0 \
-ctv q8_0
This changes KV-cache representation independently from model-weight quantization.
You can test more aggressive cache types if supported.
But do not change directly from FP16 to the smallest possible cache representation and assume the result is free.
Lower-precision KV cache has trade-offs
KV-cache quantization can reduce memory.
But possible costs include:
additional quantize/dequantize work
latency changes
numerical approximation
model-specific quality effects
For short contexts where VRAM is already sufficient, cache quantization may provide little benefit and can even hurt latency.
Use it primarily when KV-cache memory is genuinely the bottleneck.
Weight quantization and KV quantization are separate
This distinction is worth repeating.
You can have:
Q4 weights
+
FP16 KV
or:
Q4 weights
+
Q8 KV
or another supported combination.
Changing the GGUF quantization does not automatically change the KV-cache format.
This is why a heavily quantized model can still consume substantial VRAM at long context.
Fix 4: move KV cache from GPU to CPU RAM
If VRAM is extremely constrained but system RAM is available, the KV cache can sometimes be moved to CPU memory.
This saves GPU VRAM.
The trade-off is additional communication between CPU memory and GPU processing.
In current llama.cpp, GPU KV offloading is enabled by default.
To keep KV-cache buffers in CPU RAM, you can use:
--no-kv-offload
or:
-nkvo
For example:
llama-cli \
-m model.gguf \
-c 32768 \
-ngl all \
--no-kv-offload
This can make a configuration fit when GPU memory would otherwise be exhausted.
CPU KV cache can hurt performance
Moving KV state away from the GPU means the runtime must rely more heavily on CPU memory and data transfer.
The model may still work.
But prompt processing and generation can become slower.
So:
KV cache on GPU
→ more VRAM
→ generally faster
KV cache in system RAM
→ less VRAM
→ potentially slower
This is a classic memory-versus-latency trade-off.
Hugging Face cache offloading
Transformers also supports cache offloading.
The general principle is:
most KV cache
→ CPU
currently needed layer
→ GPU
The runtime transfers cache state as layers are processed.
This saves VRAM while preserving the ability to generate long sequences.
Again, the expected trade-off is lower GPU memory usage at the cost of additional transfer overhead.
Fix 5: offload fewer model layers to the GPU
If the model weights themselves do not fit, llama.cpp can split execution between:
GPU VRAM
and:
system RAM / CPU
Current llama.cpp controls GPU layer placement with:
-ngl
--n-gpu-layers
You may use:
auto
all
or an exact layer count
For example:
llama-cli \
-m model.gguf \
-ngl auto
lets the runtime choose placement.
If you deliberately need to reduce VRAM further, you can test a lower exact value.
For example:
llama-cli \
-m model.gguf \
-ngl 20
The correct number depends on model architecture.
Do not copy someone else’s layer count blindly.
Partial GPU offloading
Conceptually:
GPU:
layers 1–20
CPU / RAM:
remaining layers
This allows a model larger than available VRAM to run.
The trade-off is performance.
More work on the CPU generally means lower generation speed than a fully GPU-resident model.
Full GPU vs hybrid inference
A simplified comparison:
Full GPU
→ more VRAM required
→ usually faster
Hybrid CPU/GPU
→ less VRAM required
→ usually slower
CPU-only
→ minimal dedicated VRAM
→ often much slower
The best point depends on:
GPU
CPU
RAM bandwidth
model size
quantization
Fix 6: use a smaller model
Sometimes you have already optimized everything else.
If your hardware cannot comfortably run:
70B
the most effective fix may be:
32B
14B
8B
rather than trying increasingly extreme memory tricks.
Parameter count has a huge effect on model-weight memory.
For example, idealized 4-bit dense weights require approximately:
8B
≈ 4 GB
14B
≈ 7 GB
32B
≈ 16 GB
70B
≈ 35 GB
before quantization metadata and runtime allocations.
A smaller model can free more VRAM than almost any individual runtime setting.
Bigger is not always better for your workload
A larger model can provide better capability.
But if it forces:
heavy CPU offloading
tiny context
severe quantization
constant OOM risk
very low tokens/s
a smaller model running efficiently may provide a better user experience.
The real objective is:
quality
+
speed
+
memory
not simply the largest parameter count possible.
Fix 7: reduce parallel sequences
Single-user chat and multi-user serving have very different memory requirements.
Each active sequence requires inference state.
If you serve:
1 user
you need much less cache state than:
16 concurrent users
In llama.cpp, parallel decoding is controlled with:
-np
--parallel
For a personal workstation:
-np 1
may be enough.
If you do not need concurrency, do not allocate for it.
vLLM and concurrency
Serving engines such as vLLM dedicate substantial GPU memory to model execution and KV-cache capacity.
Current vLLM exposes controls such as:
gpu_memory_utilization
kv_cache_memory_bytes
cache_dtype
kv_offloading_size
These let administrators balance:
model memory
KV-cache capacity
concurrency
VRAM limits
A production server optimized for maximum concurrent requests should not be configured the same way as a single-user desktop.
Do not blindly maximize gpu_memory_utilization
A serving engine may allow you to reserve a high fraction of GPU memory.
That does not mean every machine should run at the maximum possible value.
Leave room for:
driver overhead
other GPU workloads
monitoring tools
multiple model instances
unexpected allocations
If the server is repeatedly near OOM, memory headroom is more useful than squeezing out one additional cache block.
Explicit KV-cache limits in vLLM
Current vLLM can automatically infer KV-cache capacity from its GPU-memory allocation.
It also exposes:
kv_cache_memory_bytes
for more direct control.
This can be useful when you want to reserve a known memory budget rather than allow cache allocation to consume whatever remains.
The right value depends on:
model size
maximum context
concurrency
GPU capacity
vLLM KV-cache dtype
vLLM also supports configuring the KV-cache data type.
Using a lower-precision cache can reduce VRAM requirements.
Again, this is independent from weight quantization.
Conceptually:
AWQ / GPTQ / FP8 weights
and:
KV-cache dtype
are separate settings.
vLLM KV-cache offloading
Current vLLM also supports CPU KV-cache offloading.
Its cache configuration exposes:
kv_offloading_size
which enables a CPU offloading buffer when configured.
This is useful when:
GPU KV capacity is insufficient
but:
system RAM is available
As with other offloading strategies, data movement can affect latency.
Fix 8: reduce batch-related memory
Prompt processing can require temporary working memory.
In llama.cpp, relevant controls include:
-b
--batch-size
and:
-ub
--ubatch-size
If you encounter OOM during large prompt processing rather than during steady generation, lowering these values can sometimes help.
For example:
llama-cli \
-m model.gguf \
-b 512 \
-ub 256
instead of much larger settings.
Do not change these unless you have evidence that batch working memory is the problem.
Smaller batches can reduce speed
Larger batches can improve prompt-processing throughput on suitable hardware.
So reducing batch size may trade:
lower peak memory
for:
slower prefill
This is another reason to diagnose the failure before changing settings.
Model loads but prompt causes OOM
This scenario strongly suggests:
weights fit
but:
context / cache / temporary memory does not
First try:
shorter context
For example:
-c 8192
instead of:
-c 32768
If that fixes the problem, model-weight quantization may not be the main issue.
OOM immediately during model loading
If the model fails before inference even begins, investigate:
model weights
GPU-layer placement
other GPU applications
runtime buffers
The strongest interventions are usually:
smaller quantization
fewer GPU layers
smaller model
Context reduction may not matter much if the weights themselves cannot be loaded.
OOM after a long conversation
If the system works initially but fails after many conversation turns, investigate:
context growth
KV-cache growth
conversation history
Start a fresh chat.
If the new conversation works, context was probably a major factor.
OOM only with multiple users
If one request works but many simultaneous requests fail, investigate:
concurrency
parallel sequence count
KV-cache capacity
batching
This is a serving-capacity problem.
Reducing model quantization may help, but the more direct solution may be reducing concurrency or cache requirements.
Close other GPU programs
Do not overlook the simplest fix.
Your GPU may already contain allocations from:
browser
desktop compositor
image generator
video generator
another LLM
PyTorch notebook
game
CUDA process
Check:
nvidia-smi
before launching the model.
Freeing:
2–4 GB
from unrelated applications may be enough to avoid OOM.
Do not use all VRAM intentionally
If your GPU has:
24 GB
do not design around:
23.99 GB
of normal usage.
You need headroom.
Model memory can vary with:
prompt length
runtime version
temporary buffers
concurrency
driver behavior
A stable configuration is better than one that survives only one carefully controlled benchmark.
How much headroom should you leave?
There is no universal percentage.
The correct amount depends on the runtime and workload.
Instead of a fixed rule, test the largest realistic workload:
longest prompt
largest expected context
maximum concurrency
normal desktop applications
Then verify that the system remains stable.
Weight memory vs KV-cache memory
A practical diagnosis is:
If VRAM is already almost full immediately after model load:
→ weight problem
If VRAM grows significantly with context:
→ KV-cache problem
If OOM appears only during prompt processing:
→ temporary/batch memory may matter
If OOM appears only under many requests:
→ concurrency/cache problem
This is much more useful than changing every setting at once.
A good llama.cpp low-VRAM baseline
Suppose a model nearly fits.
Start with something conservative:
llama-cli \
-m model.gguf \
-c 8192 \
-ngl auto \
--perf
Then monitor:
nvidia-smi
If memory remains too high, try the optimizations one at a time.
Step 1: reduce quantization
For example:
Q6_K
↓
Q5_K_M
↓
Q4_K_M
Reload and measure.
Step 2: reduce context
For example:
-c 32768
to:
-c 8192
Reload and measure.
Step 3: reduce KV precision
Test:
-ctk q8_0 -ctv q8_0
instead of the default FP16 cache.
Measure:
VRAM
quality
prompt speed
generation speed
Step 4: move KV cache to CPU
If VRAM is still insufficient:
--no-kv-offload
This can save GPU memory but may reduce performance.
Step 5: reduce GPU layers
If the weights still do not fit:
-ngl 20
or another model-appropriate count.
Allow more of the model to remain in system RAM.
Change one variable at a time
Do not simultaneously change:
quantization
context
GPU layers
KV precision
batch size
and then conclude:
it is faster now
You will not know which change mattered.
A controlled optimization workflow is:
baseline
↓
change one setting
↓
measure
↓
keep or revert
↓
change next setting
Measure both VRAM and speed
A memory optimization is not automatically a good optimization.
Suppose:
Configuration A
VRAM: 22 GB
Speed: 40 tokens/s
Configuration B
VRAM: 15 GB
Speed: 8 tokens/s
Configuration B uses less memory.
But if you already have 24 GB available, the performance loss may not be worthwhile.
The target is usually:
fit comfortably
not:
minimize VRAM at any cost
Measure quality too
Aggressive weight quantization and KV-cache quantization can affect numerical behavior.
If your workload is:
coding
reasoning
structured extraction
long-context retrieval
test real examples.
Do not optimize only against:
nvidia-smi
The model still needs to solve the task.
Optimization priority by cost
A useful mental ordering is:
Close other GPU apps
→ essentially free
Reduce unused context
→ usually low quality cost
Choose a sensible model quantization
→ moderate trade-off
Reduce KV-cache precision
→ memory/latency/quality trade-off
Move KV cache to CPU
→ speed trade-off
Offload model layers to CPU
→ potentially large speed trade-off
Choose a smaller model
→ changes model capability
The best solution is usually the least disruptive change that makes the workload fit.
Example: 24 GB GPU with a 32B model
Suppose:
GPU:
24 GB
Model:
32B Q5_K_M
Context:
32K
and you get OOM.
A sensible sequence is:
1. Reduce context to 8K
2. Test again
3. If still too large, try Q4_K_M
4. Test again
5. Reduce KV-cache precision if needed
6. Move KV cache to CPU only if necessary
Do not immediately switch to CPU-only inference.
Example: model loads but 64K context fails
Suppose:
Model loaded:
20 GB VRAM
64K prompt:
OOM
This strongly suggests that the remaining:
4 GB
was insufficient for cache and working allocations.
Possible fixes:
32K context
16K context
quantized KV cache
CPU KV cache
smaller weight quantization
The model itself may be fine.
Example: 70B on one 24 GB GPU
A conventional dense 70B model at an idealized four bits already requires roughly:
35 GB
for weights before overhead.
Therefore it cannot normally fit entirely in 24 GB VRAM.
Your choices include:
CPU/GPU hybrid inference
multiple GPUs
more aggressive quantization
smaller model
larger-memory GPU
Runtime tuning cannot make 35 GB of weights magically become 24 GB without changing representation or placement.
Multiple GPUs
Multiple GPUs can reduce the memory pressure on any single GPU by distributing model tensors and, depending on runtime, cache state.
llama.cpp supports multi-GPU split modes and tensor distribution controls.
However:
two GPUs
do not automatically behave like:
one perfectly unified pool
Interconnect bandwidth, split strategy, backend, and architecture matter.
Multi-GPU is a deployment strategy, not a free memory expansion.
Unified memory systems
On Apple Silicon, CPU and GPU share unified memory.
The problem becomes less about:
dedicated VRAM
and more about:
total available unified memory
Quantization and context still matter because model weights, KV cache, applications, and the operating system all compete for the same pool.
You still need headroom.
VRAM optimization checklist
When local inference runs out of memory, check:
Is another application using the GPU?
How large are the model weights?
What quantization am I using?
What context size is configured?
How large is the active conversation?
What KV-cache precision is being used?
Is KV cache stored on GPU?
How many layers are on the GPU?
How many sequences are active?
Are batch settings unusually large?
Answer those questions before buying hardware.
What should you change first?
For most local users, a sensible sequence is:
1. Close unnecessary GPU applications
2. Use a realistic context length
3. Choose a practical model quantization
4. Reduce KV-cache precision if long context is the issue
5. Move KV cache to CPU if VRAM is still insufficient
6. Reduce GPU layer count
7. Choose a smaller model if performance becomes unacceptable
This preserves performance as long as possible.
Bottom line
Reducing LLM VRAM usage is not one trick.
VRAM is consumed by several components:
model weights
KV cache
runtime buffers
temporary memory
parallel sequences
Different optimizations target different components.
Use:
lower-bit model quantization
to reduce weight memory.
Use:
shorter context
to reduce cache requirements.
Use:
lower-precision KV cache
to compress long-context state.
Use:
CPU KV-cache placement
when VRAM is more important than latency.
Use:
partial GPU offloading
when the model itself does not fit.
And use:
a smaller model
when the compromises required to run a larger model make the system impractical.
The objective is not to achieve the lowest possible VRAM number.
It is to find a configuration where:
the model fits
+
the context fits
+
performance remains acceptable
+
quality remains acceptable
+
there is enough memory headroom for real workloads
That is the practical way to optimize local LLM VRAM usage.