AI Hardware

How Much VRAM Do You Need for Local LLMs?

Learn how to estimate LLM VRAM requirements from model size, quantization, KV cache, context length, runtime overhead, and GPU offloading.

Approximately 17 min read

How much VRAM do you need to run a local LLM?

The tempting answer is a table such as “8B needs 8 GB” or “70B needs 48 GB.” Those numbers can be useful as rough shorthand, but they hide the variables that actually determine whether a model will run.

A better estimate is:

VRAM required
=
model weights
+ KV cache
+ activations and temporary buffers
+ runtime overhead
+ safety headroom

The biggest variable is usually the model weights. But once you use long context windows, multiple concurrent requests, or GPU-heavy serving frameworks, the other components can become large enough to turn a model that appears to fit into an out-of-memory error.

This guide shows how to estimate VRAM without pretending there is one exact number for every model size.

Quick answer

For inference, a useful first approximation for the model weights alone is:

Weight memory ≈ parameter count × bits per weight ÷ 8

That gives these idealized values:

Model size FP16/BF16 8-bit 4-bit
7B 14 GB 7 GB 3.5 GB
8B 16 GB 8 GB 4 GB
14B 28 GB 14 GB 7 GB
32B 64 GB 32 GB 16 GB
70B 140 GB 70 GB 35 GB

These are weight-only theoretical estimates in decimal GB.

They are not total VRAM requirements.

A real quantized model may consume more because of quantization metadata, scales, tensors stored at higher precision, runtime allocations, KV cache, and other buffers.

So if a calculation says a model needs exactly 23 GB and you have a 24 GB GPU, that does not mean it will safely run.

Why parameter count alone is not enough

Consider two people running the same 32B model.

User A runs:

32B model
4-bit quantization
4K context
one request

User B runs:

32B model
4-bit quantization
32K context
multiple concurrent requests

The weight memory can be similar, but the total memory requirement can be very different.

That is because model weights are only one part of inference memory.

The practical question is not:

How much VRAM does a 32B model need?

It is:

How much VRAM does this exact model, quantization, runtime, context length, and workload need?

Step 1: estimate model weight memory

Start with the model weights because they are usually the largest fixed allocation.

For a dense model, the simplest approximation is:

parameters × bytes per parameter

Common theoretical values are:

FP32       4 bytes per parameter
FP16       2 bytes per parameter
BF16       2 bytes per parameter
INT8       1 byte per parameter
4-bit      0.5 byte per parameter

So an 8B model in FP16 has an idealized weight size of:

8 billion × 2 bytes
= 16 GB

At 4 bits:

8 billion × 0.5 byte
= 4 GB

This arithmetic is useful, but quantized formats are more complicated than simply packing every weight into exactly four bits.

Quantization schemes may store:

That is why a real 4-bit model file is not guaranteed to equal exactly:

parameters × 0.5 byte

Treat the formula as a sizing baseline rather than an exact runtime measurement.

FP16, 8-bit, and 4-bit make a huge difference

Quantization is one of the most important tools for fitting larger LLMs into consumer GPUs.

For a theoretical 32B dense model:

FP16:
32B × 2 bytes
≈ 64 GB

8-bit:
32B × 1 byte
≈ 32 GB

4-bit:
32B × 0.5 byte
≈ 16 GB

That difference can move the model from multi-GPU territory to a single 24 GB GPU.

But lower weight precision is not free. The exact trade-off depends on the quantization method, model, runtime, and workload.

For an explanation of two common local deployment choices, see GGUF vs AWQ.

Step 2: account for the KV cache

After the weights, the KV cache is one of the most important sources of inference memory usage.

Autoregressive transformers repeatedly attend to previously processed tokens. Instead of recomputing the key and value tensors for every previous token at every generation step, inference engines can store them in a KV cache.

The important part for VRAM planning is:

The KV cache grows as more tokens are stored.

For many architectures, cache memory depends on factors such as:

This means a model can load successfully at startup and still run out of VRAM later when you send it a long prompt.

For a deeper explanation, see What Is KV Cache?.

Context length can change the answer

Suppose a model supports:

4K
8K
32K
128K

contexts.

That does not mean every machine that can load the model weights can run the maximum context.

Longer context normally means a larger KV cache.

So this:

Model fits in 18 GB
GPU has 24 GB

does not automatically imply:

6 GB is plenty for every context length

The remaining memory also needs to cover runtime allocations, temporary buffers, and the cache.

If you are close to the VRAM limit, reducing context length can be one of the most effective ways to avoid an OOM error.

Architecture matters: MHA, GQA, and MQA

Two models with the same parameter count and context length can have different KV cache requirements.

One reason is the attention architecture.

Traditional multi-head attention can maintain separate key and value heads corresponding to many query heads.

Architectures using Grouped-Query Attention (GQA) or Multi-Query Attention (MQA) share key/value heads across multiple query heads.

That can substantially reduce KV cache memory.

Therefore:

Model A: 32B
Model B: 32B

does not imply that both need the same cache size.

This is another reason universal “VRAM by parameter count” charts should be treated as approximations.

Step 3: leave memory for runtime overhead

The model does not own every byte of GPU memory.

Your inference stack may allocate memory for:

The exact amount depends heavily on the runtime.

For this reason, trying to fill 100% of advertised VRAM with model weights is usually a bad sizing strategy.

A GPU marketed as 24 GB is not equivalent to:

24 GB available exclusively for model weights.

Why vLLM may use more VRAM than expected

vLLM is designed for high-throughput inference and actively manages GPU memory.

Its configuration includes a gpu_memory_utilization setting that controls the fraction of GPU memory available to the model executor.

vLLM also allocates memory for its KV cache.

That makes the following comparison misleading:

Model uses 17 GB in one local runtime
therefore
it must use 17 GB in vLLM

Different inference engines make different memory/performance trade-offs.

A local interactive runtime and a high-throughput serving runtime should not be expected to have identical VRAM behavior.

With vLLM, concurrency matters especially because KV cache capacity directly affects how many tokens and requests can be handled simultaneously.

Concurrency matters

For one person chatting with a model locally, you may have only one active sequence.

For a server, you might have:

User 1: 8,000 tokens
User 2: 12,000 tokens
User 3: 4,000 tokens
User 4: 20,000 tokens

at the same time.

The model weights are still loaded once, but the serving system needs cache capacity for active requests.

So:

single-user VRAM requirement

and:

production serving VRAM requirement

can be very different even when the underlying model is identical.

This is one reason a 24 GB GPU may feel spacious for local experimentation but restrictive for high-concurrency serving.

How much VRAM does an 8B model need?

For an 8B dense model, theoretical weight memory is approximately:

FP16/BF16: 16 GB
8-bit:      8 GB
4-bit:      4 GB

In practice, total VRAM will be higher.

An 8B model is therefore a comfortable class for many modern consumer GPUs when quantized.

A 4-bit 8B model can often leave substantial room for context and runtime overhead on GPUs with 12 GB, 16 GB, or more VRAM.

The exact maximum context still depends on architecture and runtime.

How much VRAM does a 14B model need?

Idealized weights:

FP16/BF16: 28 GB
8-bit:     14 GB
4-bit:      7 GB

A 14B model is an important transition point.

FP16 generally exceeds the VRAM capacity of common 24 GB consumer cards, while quantization makes single-GPU use much more practical.

For local inference, 12 GB to 16 GB GPUs may be usable with sufficiently compact quantization and sensible context settings.

A 24 GB GPU provides much more headroom.

How much VRAM does a 32B model need?

Idealized weights:

FP16/BF16: 64 GB
8-bit:     32 GB
4-bit:     16 GB

This is where a 24 GB GPU becomes especially useful.

A roughly 4-bit 32B model can have enough room to fit its weights on a 24 GB GPU while leaving some memory for the KV cache and runtime.

But “32B 4-bit fits on 24 GB” should not be interpreted as a guarantee for:

maximum context
+
large batch size
+
high concurrency

all at once.

If your workload is long-context or server-oriented, the remaining VRAM becomes important quickly.

How much VRAM does a 70B model need?

Idealized weight memory:

FP16/BF16: 140 GB
8-bit:      70 GB
4-bit:      35 GB

Even the idealized 4-bit weight size is greater than 24 GB.

Therefore, a dense 70B model generally cannot have all of its 4-bit weights resident on a single 24 GB GPU.

Your options include:

This is where CPU+GPU hybrid inference becomes useful.

Can a 24 GB GPU run a 70B model?

Yes, but the meaning of “run” matters.

If you mean:

Can every layer and all supporting inference state remain entirely inside 24 GB VRAM?

Normally not for a typical 4-bit dense 70B model.

If you mean:

Can I generate text from a 70B-class quantized model using a 24 GB GPU plus system RAM?

That can be possible with runtimes that support partial GPU offloading.

llama.cpp, for example, supports CPU+GPU hybrid inference specifically so models larger than available VRAM can still be used.

The trade-off is performance.

Weights or operations that remain on the CPU depend on system memory bandwidth and CPU performance rather than the much higher bandwidth available entirely on the GPU.

So “it runs” and “it runs fast” are separate questions.

What about 12 GB VRAM?

A 12 GB GPU can still be highly useful for local LLMs.

Good targets commonly include quantized models in classes such as:

7B
8B
12B
14B

depending on the exact quantization and context requirements.

A theoretical 14B model at 4 bits uses about 7 GB for weights, leaving some capacity for the rest of inference.

Larger models may still be usable through partial GPU offloading, but the proportion remaining in system RAM grows.

For many users, 12 GB is enough to learn local AI and run capable models without needing expensive workstation hardware.

What about 16 GB VRAM?

16 GB offers noticeably more flexibility.

It can accommodate:

A theoretical 32B model at exactly 4 bits already requires around 16 GB just for weights, however.

That means a 16 GB GPU should not be treated as a guaranteed full-GPU home for every 32B 4-bit model.

There still needs to be memory for everything else.

What about 24 GB VRAM?

24 GB is a particularly useful capacity for enthusiast local AI.

It provides enough memory to comfortably run many quantized models below the 30B range and can make many 30B-class models practical entirely on GPU.

A common planning target is:

4-bit 32B weights ≈ 16 GB theoretical

which leaves meaningful headroom on a 24 GB card.

Actual usage still depends on the quantization format, architecture, runtime, and context length.

A 24 GB GPU does not magically make every 32B workload fit, but it gives far more flexibility than 12 GB or 16 GB.

What about 48 GB VRAM?

48 GB moves into a substantially different class.

A theoretical 4-bit 70B dense model requires around:

35 GB

of weight memory.

That means some 70B-class quantized models can fit fully on a 48 GB GPU with room remaining for inference state.

Whether the available remainder is sufficient for your desired context and concurrency still requires testing.

For serious long-context work or serving, even 48 GB can become constrained by the KV cache.

Model file size is a useful shortcut for GGUF

When using GGUF models, the downloaded file size gives a useful first approximation of the memory needed to store the model tensors.

For example, if your GGUF file is:

18.5 GB

you already know the model cannot use only:

12 GB VRAM

if the goal is to keep all of those weights on the GPU.

But file size is still not total runtime VRAM.

You need additional memory for cache and computation.

So a useful mental model is:

GGUF file size
≈ starting point for weight memory

not

GGUF file size
= total VRAM requirement

GPU offloading changes the equation

A model does not always have to fit entirely into VRAM.

llama.cpp can split work between CPU and GPU.

Conceptually:

GPU VRAM:
some model layers
+ GPU-side cache/buffers

System RAM:
remaining model layers
+ CPU-side data

This lets you use models larger than your GPU.

For example:

24 GB GPU
64 GB or 128 GB system RAM

can run configurations that would be impossible if every model weight had to stay in VRAM.

The cost is usually lower inference speed compared with full GPU residency.

How severe the slowdown is depends on the amount offloaded, CPU, RAM bandwidth, GPU, model architecture, and runtime.

System RAM still matters

Local LLM discussions often focus only on VRAM, but system RAM becomes important when:

A machine with a powerful GPU but very limited system RAM can still be awkward for large local models.

For users interested in experimenting beyond the GPU’s native capacity, system RAM provides valuable flexibility.

Unified-memory systems are different

Apple Silicon does not divide memory into conventional discrete system RAM and GPU VRAM in the same way as a desktop with an NVIDIA graphics card.

CPU and GPU operate from a unified memory pool.

That means a Mac advertised with:

64 GB unified memory

should not be compared directly with:

64 GB system RAM
+
24 GB NVIDIA VRAM

The memory architecture and bandwidth characteristics are different.

When comparing local AI hardware, look at the entire platform rather than comparing memory capacity labels alone.

Don’t confuse inference memory with training memory

This article is about inference.

Training or full fine-tuning requires much more memory because the system may need to store:

That can multiply memory requirements far beyond simple weight storage.

Techniques such as LoRA and QLoRA reduce training memory requirements, but a statement like:

“This model fits in 12 GB for inference”

does not imply:

“I can fully train this model in 12 GB.”

They are different workloads.

A better way to estimate VRAM before downloading

Use this process.

1. Choose the exact model

Do not start with:

I want a 32B model.

Start with the exact repository or file.

Different architectures and model variants differ.

2. Choose the exact quantization

For example:

FP16
BF16
INT8
AWQ 4-bit
GGUF Q4_K_M
GGUF Q5_K_M

A parameter count without precision is not enough.

3. Estimate weight memory

Use:

parameter count × bits ÷ 8

as a theoretical baseline.

For GGUF, the actual file size is often an even more practical starting point.

4. Decide your context requirement

Ask whether you really need:

128K

or whether:

8K
16K
32K

covers your workload.

Maximum advertised context is not automatically the best configuration.

5. Decide whether this is single-user or serving

One local chat session has very different cache requirements from many simultaneous API requests.

6. Leave headroom

Never plan around using exactly 100% of the GPU’s advertised memory.

7. Test the real workload

The best measurement is still the actual model running in the actual runtime with the context and concurrency you intend to use.

Monitor peak memory, not just memory immediately after model loading.

A practical GPU selection guide

Instead of treating these as guarantees, use them as broad planning categories.

VRAM Practical local-AI positioning
8 GB Smaller quantized models; context and offload need attention
12 GB Strong starting point for 7B/8B and many quantized mid-size models
16 GB More room for larger quantizations and context; some larger models become practical
24 GB Very flexible enthusiast tier; many 30B-class quantized models become viable
32 GB More headroom for larger models, context, and serving
48 GB Many quantized 70B-class configurations become possible fully on GPU
80 GB+ Large-model inference and much more generous serving capacity

These categories deliberately avoid claiming that every model of a given parameter count will fit.

The exact model still wins over the chart.

Example: choosing between a 16 GB and 24 GB GPU

Suppose your target is a 32B model at roughly 4-bit weight precision.

The theoretical weight requirement is:

32B × 4 bits ÷ 8
≈ 16 GB

On a 16 GB GPU, the model weights alone approximately consume the entire theoretical capacity.

That leaves effectively no planned room for:

So this is a poor full-GPU sizing target.

On a 24 GB GPU:

24 GB
- ~16 GB theoretical weights
= ~8 GB before other overhead

That does not guarantee success, but the configuration is much more realistic.

This illustrates why buying a GPU based only on the minimum model weight calculation can be frustrating.

Example: why a model loads but crashes on a long prompt

Imagine:

GPU: 24 GB
Model + runtime after loading: 20 GB

You send a short prompt and generation works.

Then you submit a much longer document.

The KV cache grows as more tokens are retained.

Eventually the process attempts another GPU allocation and fails.

The model itself did not suddenly become larger.

The working inference state became larger.

Possible fixes include:

This is why “the model loads” is not the same as “the workload fits.”

Can you reduce KV cache memory?

Depending on the runtime, yes.

Possible techniques include:

Hugging Face Transformers, for example, provides both offloaded and quantized cache implementations.

These features trade memory against other factors such as latency and compute overhead.

If the normal cache already fits comfortably, compressing or offloading it may not make the workload faster.

Should you buy more VRAM or use heavier quantization?

It depends on what you value.

More VRAM lets you use:

More aggressive quantization lets you fit larger models on cheaper hardware.

Neither is universally better.

A smaller model at higher precision can sometimes be a better choice than a much larger model at extremely aggressive quantization, depending on the task.

Choose based on measured usefulness, not parameter count alone.

The most useful rule

When evaluating whether a model will fit, do not ask:

How much VRAM does a 7B, 14B, 32B, or 70B model need?

Ask:

What is the weight memory of this exact quantization, and how much additional memory will my context, KV cache, runtime, and concurrency require?

That question produces much better hardware decisions.

Bottom line

A simple parameter-count table is useful for estimating weight memory, but it cannot tell you the full VRAM requirement.

Remember:

Total VRAM
=
weights
+ KV cache
+ runtime buffers
+ temporary allocations
+ headroom

For weight memory alone:

FP16/BF16 ≈ 2 bytes per parameter
INT8      ≈ 1 byte per parameter
4-bit     ≈ 0.5 byte per parameter

Then add room for the rest of inference.

For local users, quantization and CPU/GPU offloading can make surprisingly large models usable on consumer hardware.

For serving systems, context length and concurrency make KV cache capacity increasingly important.

And if you are choosing hardware, do not buy the smallest GPU that satisfies the theoretical weight calculation.

Buy enough headroom for the workload you actually intend to run.

Sources and further reading

Continue reading