How Much VRAM Do You Need for Local LLMs?
Learn how to estimate LLM VRAM requirements from model size, quantization, KV cache, context length, runtime overhead, and GPU offloading.
Approximately 17 min read
How much VRAM do you need to run a local LLM?
The tempting answer is a table such as “8B needs 8 GB” or “70B needs 48 GB.” Those numbers can be useful as rough shorthand, but they hide the variables that actually determine whether a model will run.
A better estimate is:
VRAM required
=
model weights
+ KV cache
+ activations and temporary buffers
+ runtime overhead
+ safety headroom
The biggest variable is usually the model weights. But once you use long context windows, multiple concurrent requests, or GPU-heavy serving frameworks, the other components can become large enough to turn a model that appears to fit into an out-of-memory error.
This guide shows how to estimate VRAM without pretending there is one exact number for every model size.
Quick answer
For inference, a useful first approximation for the model weights alone is:
Weight memory ≈ parameter count × bits per weight ÷ 8
That gives these idealized values:
| Model size | FP16/BF16 | 8-bit | 4-bit |
|---|---|---|---|
| 7B | 14 GB | 7 GB | 3.5 GB |
| 8B | 16 GB | 8 GB | 4 GB |
| 14B | 28 GB | 14 GB | 7 GB |
| 32B | 64 GB | 32 GB | 16 GB |
| 70B | 140 GB | 70 GB | 35 GB |
These are weight-only theoretical estimates in decimal GB.
They are not total VRAM requirements.
A real quantized model may consume more because of quantization metadata, scales, tensors stored at higher precision, runtime allocations, KV cache, and other buffers.
So if a calculation says a model needs exactly 23 GB and you have a 24 GB GPU, that does not mean it will safely run.
Why parameter count alone is not enough
Consider two people running the same 32B model.
User A runs:
32B model
4-bit quantization
4K context
one request
User B runs:
32B model
4-bit quantization
32K context
multiple concurrent requests
The weight memory can be similar, but the total memory requirement can be very different.
That is because model weights are only one part of inference memory.
The practical question is not:
How much VRAM does a 32B model need?
It is:
How much VRAM does this exact model, quantization, runtime, context length, and workload need?
Step 1: estimate model weight memory
Start with the model weights because they are usually the largest fixed allocation.
For a dense model, the simplest approximation is:
parameters × bytes per parameter
Common theoretical values are:
FP32 4 bytes per parameter
FP16 2 bytes per parameter
BF16 2 bytes per parameter
INT8 1 byte per parameter
4-bit 0.5 byte per parameter
So an 8B model in FP16 has an idealized weight size of:
8 billion × 2 bytes
= 16 GB
At 4 bits:
8 billion × 0.5 byte
= 4 GB
This arithmetic is useful, but quantized formats are more complicated than simply packing every weight into exactly four bits.
Quantization schemes may store:
- scaling values
- zero points
- group metadata
- some tensors at higher precision
- embeddings or output layers differently
- format metadata
That is why a real 4-bit model file is not guaranteed to equal exactly:
parameters × 0.5 byte
Treat the formula as a sizing baseline rather than an exact runtime measurement.
FP16, 8-bit, and 4-bit make a huge difference
Quantization is one of the most important tools for fitting larger LLMs into consumer GPUs.
For a theoretical 32B dense model:
FP16:
32B × 2 bytes
≈ 64 GB
8-bit:
32B × 1 byte
≈ 32 GB
4-bit:
32B × 0.5 byte
≈ 16 GB
That difference can move the model from multi-GPU territory to a single 24 GB GPU.
But lower weight precision is not free. The exact trade-off depends on the quantization method, model, runtime, and workload.
For an explanation of two common local deployment choices, see GGUF vs AWQ.
Step 2: account for the KV cache
After the weights, the KV cache is one of the most important sources of inference memory usage.
Autoregressive transformers repeatedly attend to previously processed tokens. Instead of recomputing the key and value tensors for every previous token at every generation step, inference engines can store them in a KV cache.
The important part for VRAM planning is:
The KV cache grows as more tokens are stored.
For many architectures, cache memory depends on factors such as:
- number of transformer layers
- number of KV heads
- head dimension
- cache data type
- context length
- number of active sequences
- batch size or concurrency
This means a model can load successfully at startup and still run out of VRAM later when you send it a long prompt.
For a deeper explanation, see What Is KV Cache?.
Context length can change the answer
Suppose a model supports:
4K
8K
32K
128K
contexts.
That does not mean every machine that can load the model weights can run the maximum context.
Longer context normally means a larger KV cache.
So this:
Model fits in 18 GB
GPU has 24 GB
does not automatically imply:
6 GB is plenty for every context length
The remaining memory also needs to cover runtime allocations, temporary buffers, and the cache.
If you are close to the VRAM limit, reducing context length can be one of the most effective ways to avoid an OOM error.
Architecture matters: MHA, GQA, and MQA
Two models with the same parameter count and context length can have different KV cache requirements.
One reason is the attention architecture.
Traditional multi-head attention can maintain separate key and value heads corresponding to many query heads.
Architectures using Grouped-Query Attention (GQA) or Multi-Query Attention (MQA) share key/value heads across multiple query heads.
That can substantially reduce KV cache memory.
Therefore:
Model A: 32B
Model B: 32B
does not imply that both need the same cache size.
This is another reason universal “VRAM by parameter count” charts should be treated as approximations.
Step 3: leave memory for runtime overhead
The model does not own every byte of GPU memory.
Your inference stack may allocate memory for:
- CUDA context
- kernels
- intermediate activations
- temporary tensors
- compute workspaces
- graph capture
- memory pools
- tokenizer or multimodal components
- output buffers
- quantization-related workspaces
The exact amount depends heavily on the runtime.
For this reason, trying to fill 100% of advertised VRAM with model weights is usually a bad sizing strategy.
A GPU marketed as 24 GB is not equivalent to:
24 GB available exclusively for model weights.
Why vLLM may use more VRAM than expected
vLLM is designed for high-throughput inference and actively manages GPU memory.
Its configuration includes a gpu_memory_utilization setting that controls the fraction of GPU memory available to the model executor.
vLLM also allocates memory for its KV cache.
That makes the following comparison misleading:
Model uses 17 GB in one local runtime
therefore
it must use 17 GB in vLLM
Different inference engines make different memory/performance trade-offs.
A local interactive runtime and a high-throughput serving runtime should not be expected to have identical VRAM behavior.
With vLLM, concurrency matters especially because KV cache capacity directly affects how many tokens and requests can be handled simultaneously.
Concurrency matters
For one person chatting with a model locally, you may have only one active sequence.
For a server, you might have:
User 1: 8,000 tokens
User 2: 12,000 tokens
User 3: 4,000 tokens
User 4: 20,000 tokens
at the same time.
The model weights are still loaded once, but the serving system needs cache capacity for active requests.
So:
single-user VRAM requirement
and:
production serving VRAM requirement
can be very different even when the underlying model is identical.
This is one reason a 24 GB GPU may feel spacious for local experimentation but restrictive for high-concurrency serving.
How much VRAM does an 8B model need?
For an 8B dense model, theoretical weight memory is approximately:
FP16/BF16: 16 GB
8-bit: 8 GB
4-bit: 4 GB
In practice, total VRAM will be higher.
An 8B model is therefore a comfortable class for many modern consumer GPUs when quantized.
A 4-bit 8B model can often leave substantial room for context and runtime overhead on GPUs with 12 GB, 16 GB, or more VRAM.
The exact maximum context still depends on architecture and runtime.
How much VRAM does a 14B model need?
Idealized weights:
FP16/BF16: 28 GB
8-bit: 14 GB
4-bit: 7 GB
A 14B model is an important transition point.
FP16 generally exceeds the VRAM capacity of common 24 GB consumer cards, while quantization makes single-GPU use much more practical.
For local inference, 12 GB to 16 GB GPUs may be usable with sufficiently compact quantization and sensible context settings.
A 24 GB GPU provides much more headroom.
How much VRAM does a 32B model need?
Idealized weights:
FP16/BF16: 64 GB
8-bit: 32 GB
4-bit: 16 GB
This is where a 24 GB GPU becomes especially useful.
A roughly 4-bit 32B model can have enough room to fit its weights on a 24 GB GPU while leaving some memory for the KV cache and runtime.
But “32B 4-bit fits on 24 GB” should not be interpreted as a guarantee for:
maximum context
+
large batch size
+
high concurrency
all at once.
If your workload is long-context or server-oriented, the remaining VRAM becomes important quickly.
How much VRAM does a 70B model need?
Idealized weight memory:
FP16/BF16: 140 GB
8-bit: 70 GB
4-bit: 35 GB
Even the idealized 4-bit weight size is greater than 24 GB.
Therefore, a dense 70B model generally cannot have all of its 4-bit weights resident on a single 24 GB GPU.
Your options include:
- multiple GPUs
- CPU/GPU offloading
- a more aggressive quantization
- a runtime that supports split execution
- choosing a smaller model
This is where CPU+GPU hybrid inference becomes useful.
Can a 24 GB GPU run a 70B model?
Yes, but the meaning of “run” matters.
If you mean:
Can every layer and all supporting inference state remain entirely inside 24 GB VRAM?
Normally not for a typical 4-bit dense 70B model.
If you mean:
Can I generate text from a 70B-class quantized model using a 24 GB GPU plus system RAM?
That can be possible with runtimes that support partial GPU offloading.
llama.cpp, for example, supports CPU+GPU hybrid inference specifically so models larger than available VRAM can still be used.
The trade-off is performance.
Weights or operations that remain on the CPU depend on system memory bandwidth and CPU performance rather than the much higher bandwidth available entirely on the GPU.
So “it runs” and “it runs fast” are separate questions.
What about 12 GB VRAM?
A 12 GB GPU can still be highly useful for local LLMs.
Good targets commonly include quantized models in classes such as:
7B
8B
12B
14B
depending on the exact quantization and context requirements.
A theoretical 14B model at 4 bits uses about 7 GB for weights, leaving some capacity for the rest of inference.
Larger models may still be usable through partial GPU offloading, but the proportion remaining in system RAM grows.
For many users, 12 GB is enough to learn local AI and run capable models without needing expensive workstation hardware.
What about 16 GB VRAM?
16 GB offers noticeably more flexibility.
It can accommodate:
- higher-quality quantizations of smaller models
- larger context windows
- many 14B-class configurations
- some larger models with aggressive quantization or partial offload
A theoretical 32B model at exactly 4 bits already requires around 16 GB just for weights, however.
That means a 16 GB GPU should not be treated as a guaranteed full-GPU home for every 32B 4-bit model.
There still needs to be memory for everything else.
What about 24 GB VRAM?
24 GB is a particularly useful capacity for enthusiast local AI.
It provides enough memory to comfortably run many quantized models below the 30B range and can make many 30B-class models practical entirely on GPU.
A common planning target is:
4-bit 32B weights ≈ 16 GB theoretical
which leaves meaningful headroom on a 24 GB card.
Actual usage still depends on the quantization format, architecture, runtime, and context length.
A 24 GB GPU does not magically make every 32B workload fit, but it gives far more flexibility than 12 GB or 16 GB.
What about 48 GB VRAM?
48 GB moves into a substantially different class.
A theoretical 4-bit 70B dense model requires around:
35 GB
of weight memory.
That means some 70B-class quantized models can fit fully on a 48 GB GPU with room remaining for inference state.
Whether the available remainder is sufficient for your desired context and concurrency still requires testing.
For serious long-context work or serving, even 48 GB can become constrained by the KV cache.
Model file size is a useful shortcut for GGUF
When using GGUF models, the downloaded file size gives a useful first approximation of the memory needed to store the model tensors.
For example, if your GGUF file is:
18.5 GB
you already know the model cannot use only:
12 GB VRAM
if the goal is to keep all of those weights on the GPU.
But file size is still not total runtime VRAM.
You need additional memory for cache and computation.
So a useful mental model is:
GGUF file size
≈ starting point for weight memory
not
GGUF file size
= total VRAM requirement
GPU offloading changes the equation
A model does not always have to fit entirely into VRAM.
llama.cpp can split work between CPU and GPU.
Conceptually:
GPU VRAM:
some model layers
+ GPU-side cache/buffers
System RAM:
remaining model layers
+ CPU-side data
This lets you use models larger than your GPU.
For example:
24 GB GPU
64 GB or 128 GB system RAM
can run configurations that would be impossible if every model weight had to stay in VRAM.
The cost is usually lower inference speed compared with full GPU residency.
How severe the slowdown is depends on the amount offloaded, CPU, RAM bandwidth, GPU, model architecture, and runtime.
System RAM still matters
Local LLM discussions often focus only on VRAM, but system RAM becomes important when:
- loading large model files
- using CPU-only inference
- using CPU/GPU hybrid inference
- running several models
- using memory-mapped files
- doing other work while inference is active
A machine with a powerful GPU but very limited system RAM can still be awkward for large local models.
For users interested in experimenting beyond the GPU’s native capacity, system RAM provides valuable flexibility.
Unified-memory systems are different
Apple Silicon does not divide memory into conventional discrete system RAM and GPU VRAM in the same way as a desktop with an NVIDIA graphics card.
CPU and GPU operate from a unified memory pool.
That means a Mac advertised with:
64 GB unified memory
should not be compared directly with:
64 GB system RAM
+
24 GB NVIDIA VRAM
The memory architecture and bandwidth characteristics are different.
When comparing local AI hardware, look at the entire platform rather than comparing memory capacity labels alone.
Don’t confuse inference memory with training memory
This article is about inference.
Training or full fine-tuning requires much more memory because the system may need to store:
- model weights
- gradients
- optimizer states
- activations for backpropagation
That can multiply memory requirements far beyond simple weight storage.
Techniques such as LoRA and QLoRA reduce training memory requirements, but a statement like:
“This model fits in 12 GB for inference”
does not imply:
“I can fully train this model in 12 GB.”
They are different workloads.
A better way to estimate VRAM before downloading
Use this process.
1. Choose the exact model
Do not start with:
I want a 32B model.
Start with the exact repository or file.
Different architectures and model variants differ.
2. Choose the exact quantization
For example:
FP16
BF16
INT8
AWQ 4-bit
GGUF Q4_K_M
GGUF Q5_K_M
A parameter count without precision is not enough.
3. Estimate weight memory
Use:
parameter count × bits ÷ 8
as a theoretical baseline.
For GGUF, the actual file size is often an even more practical starting point.
4. Decide your context requirement
Ask whether you really need:
128K
or whether:
8K
16K
32K
covers your workload.
Maximum advertised context is not automatically the best configuration.
5. Decide whether this is single-user or serving
One local chat session has very different cache requirements from many simultaneous API requests.
6. Leave headroom
Never plan around using exactly 100% of the GPU’s advertised memory.
7. Test the real workload
The best measurement is still the actual model running in the actual runtime with the context and concurrency you intend to use.
Monitor peak memory, not just memory immediately after model loading.
A practical GPU selection guide
Instead of treating these as guarantees, use them as broad planning categories.
| VRAM | Practical local-AI positioning |
|---|---|
| 8 GB | Smaller quantized models; context and offload need attention |
| 12 GB | Strong starting point for 7B/8B and many quantized mid-size models |
| 16 GB | More room for larger quantizations and context; some larger models become practical |
| 24 GB | Very flexible enthusiast tier; many 30B-class quantized models become viable |
| 32 GB | More headroom for larger models, context, and serving |
| 48 GB | Many quantized 70B-class configurations become possible fully on GPU |
| 80 GB+ | Large-model inference and much more generous serving capacity |
These categories deliberately avoid claiming that every model of a given parameter count will fit.
The exact model still wins over the chart.
Example: choosing between a 16 GB and 24 GB GPU
Suppose your target is a 32B model at roughly 4-bit weight precision.
The theoretical weight requirement is:
32B × 4 bits ÷ 8
≈ 16 GB
On a 16 GB GPU, the model weights alone approximately consume the entire theoretical capacity.
That leaves effectively no planned room for:
- KV cache
- runtime overhead
- temporary allocations
So this is a poor full-GPU sizing target.
On a 24 GB GPU:
24 GB
- ~16 GB theoretical weights
= ~8 GB before other overhead
That does not guarantee success, but the configuration is much more realistic.
This illustrates why buying a GPU based only on the minimum model weight calculation can be frustrating.
Example: why a model loads but crashes on a long prompt
Imagine:
GPU: 24 GB
Model + runtime after loading: 20 GB
You send a short prompt and generation works.
Then you submit a much longer document.
The KV cache grows as more tokens are retained.
Eventually the process attempts another GPU allocation and fails.
The model itself did not suddenly become larger.
The working inference state became larger.
Possible fixes include:
- shorter context
- fewer concurrent requests
- lower KV cache precision when supported
- cache offloading when supported
- smaller model quantization
- partial model offloading
- a GPU with more memory
This is why “the model loads” is not the same as “the workload fits.”
Can you reduce KV cache memory?
Depending on the runtime, yes.
Possible techniques include:
- lower-precision KV cache
- KV cache offloading
- shorter context
- lower concurrency
- architectures using GQA or MQA
Hugging Face Transformers, for example, provides both offloaded and quantized cache implementations.
These features trade memory against other factors such as latency and compute overhead.
If the normal cache already fits comfortably, compressing or offloading it may not make the workload faster.
Should you buy more VRAM or use heavier quantization?
It depends on what you value.
More VRAM lets you use:
- larger models
- higher-quality quantization
- longer context
- greater concurrency
- fewer CPU/GPU transfers
More aggressive quantization lets you fit larger models on cheaper hardware.
Neither is universally better.
A smaller model at higher precision can sometimes be a better choice than a much larger model at extremely aggressive quantization, depending on the task.
Choose based on measured usefulness, not parameter count alone.
The most useful rule
When evaluating whether a model will fit, do not ask:
How much VRAM does a 7B, 14B, 32B, or 70B model need?
Ask:
What is the weight memory of this exact quantization, and how much additional memory will my context, KV cache, runtime, and concurrency require?
That question produces much better hardware decisions.
Bottom line
A simple parameter-count table is useful for estimating weight memory, but it cannot tell you the full VRAM requirement.
Remember:
Total VRAM
=
weights
+ KV cache
+ runtime buffers
+ temporary allocations
+ headroom
For weight memory alone:
FP16/BF16 ≈ 2 bytes per parameter
INT8 ≈ 1 byte per parameter
4-bit ≈ 0.5 byte per parameter
Then add room for the rest of inference.
For local users, quantization and CPU/GPU offloading can make surprisingly large models usable on consumer hardware.
For serving systems, context length and concurrency make KV cache capacity increasingly important.
And if you are choosing hardware, do not buy the smallest GPU that satisfies the theoretical weight calculation.
Buy enough headroom for the workload you actually intend to run.