Troubleshooting

llama.cpp 'failed to allocate buffer for kv cache': How to Fix It

Fix llama.cpp KV-cache allocation failures by reducing context or concurrency, checking cache types and device placement, and reading the backend allocator error.

Approximately 9 min read

If llama.cpp ends context initialization with:

failed to allocate buffer for kv cache

use this order:

  1. Reduce context length.
  2. Reduce concurrency / parallel slots.
  3. Inspect the K and V cache datatypes.
  4. Leave VRAM headroom instead of maximizing model offload.
  5. Check which GPU or backend receives the failed allocation.
  6. Read the exact allocator failure immediately before the final KV-cache error.
  7. If the memory should fit, investigate a backend-specific limit or regression.

The central mistake is assuming model weights fit = the runtime fits. It does not.

llama.cpp allocates model weights and runtime state separately. A GGUF can load successfully and then fail while allocating KV cache, compute buffers, or another backend buffer required to create the context.

What the error actually means

Current llama.cpp KV-cache code builds cache tensors for one or more backend buffer types and then requests a real backend allocation. If that allocation returns no buffer, the runtime throws:

failed to allocate buffer for kv cache

That final line is intentionally generic. The more useful line is often immediately above it:

CUDA... cudaMalloc failed: out of memory
Vulkan... ErrorOutOfDeviceMemory
ROCm... failed to allocate ... on device 1
CPU... insufficient memory
OpenVINO... allocation limit / backend error

The final exception tells you which subsystem failed. The preceding backend message often tells you why and where.

Do not paste only the last line into a search box and discard the allocator message above it.

Why model weights fitting is not enough

A local-LLM memory budget can contain several independent pieces:

model weights
+ KV cache
+ compute / graph buffers
+ output buffers
+ backend workspaces
+ driver/runtime overhead
+ other processes using the same device

A 24 GiB GPU with a model using most of that VRAM does not have 24 GiB available for the KV cache.

Likewise, moving enough weights to a GPU to make the model “fully offloaded” can leave too little contiguous device memory for the context you asked llama.cpp to create.

This is why reducing --gpu-layers can sometimes make a larger context possible: the model runs with fewer weights resident on that device, leaving headroom for runtime allocations.

1. Reduce context first

Context length is one of the strongest KV-memory controls.

Start with a clearly smaller value and verify that the error boundary moves:

llama-cli   -m model.gguf   --ctx-size 8192   -n 1   -p "test"

If 8192 works, test upward deliberately rather than jumping back to the model’s advertised maximum.

An advertised 128K, 256K, or 1M context means the architecture or checkpoint is designed to address that range under suitable conditions. It does not guarantee that your hardware, KV precision, backend, concurrency and offload layout can allocate that context.

Upstream issue #8101 is an extreme illustration: a 1,048,576-token context request led to a huge CPU KV allocation request and failure. The useful lesson is not the exact number from that machine. It is that context settings can dominate runtime memory even when model loading itself succeeded.

For the underlying concept, see What Is KV Cache? and What Does Context Length Mean?.

2. Reduce concurrency / parallel slots

Server concurrency can change memory requirements and cache topology.

In current llama.cpp, the common parameter path maps the requested parallel count into the context’s maximum sequence count. Some architectures and cache modes maintain per-sequence or per-stream state, so concurrency is not always free.

For a diagnostic run:

llama-server   -m model.gguf   --ctx-size 32768   --parallel 1

Then compare against your normal parallel setting.

Upstream issue #26654 is a useful architecture-specific example: a DeepSeek-V4 configuration failed KV allocation under the default parallel behavior and started with --parallel 1. That does not mean every model’s KV cache is a simple “parallel slots times context” formula. It means parallelism is a real variable and should be isolated.

If reducing parallelism changes the failure, keep it in the bug report.

3. Inspect KV cache datatype

Current llama.cpp exposes separate K and V cache types:

--cache-type-k
--cache-type-v

The current CLI documents f16 as the default and lists lower-bit cache types such as q8_0 and q4 variants where supported.

Lower-precision cache types can reduce KV memory, but this is not a universal compatibility switch. Architecture and backend support matter, and current llama.cpp requires Flash Attention for some quantized-V configurations.

Check the exact options your binary supports:

llama-cli --help | grep -E 'cache-type|flash-attn'

Then change one variable at a time.

For example, if your workload is currently using f32 K/V because of a wrapper default or copied configuration, moving to a supported lower-precision type can materially change the cache budget. Do not assume the best choice from a forum post written for another architecture.

4. Leave VRAM headroom instead of maximizing model offload

If you force nearly every possible weight onto the GPU, the model may fit while the context does not.

Current llama.cpp exposes:

-ngl, --gpu-layers, --n-gpu-layers

Use a smaller GPU-layer count as a diagnostic. The goal is not necessarily to keep that configuration forever; it is to test whether moving some model weights out of VRAM creates enough room for KV and compute allocations.

The sequence is:

fully/mostly offloaded -> KV allocation fails
reduce GPU-resident weights -> retry same context
works -> VRAM headroom was part of the failure

Also remember that lowering context and lowering GPU layers trade different resources. Find a combination that matches the workload rather than optimizing a single headline number.

See How to Reduce LLM VRAM Usage for the broader memory-budget view.

5. On multi-GPU systems, inspect the device that actually failed

Do not add the VRAM of every GPU and conclude that an individual allocation must fit.

With the default layer split, current llama.cpp documents that each GPU owns a contiguous slice of model layers and that the KV cache for a layer lives on the GPU that owns that layer. Experimental tensor split behaves differently and can split both weights and KV.

That means:

GPU 0: nearly full
GPU 1: several GiB free
aggregate free VRAM: looks adequate
allocation target: GPU 0
result: failure

is entirely possible.

Start with:

llama-cli --list-devices

On NVIDIA:

nvidia-smi

Then read the allocator line for the device number.

An upstream ROCm report, issue #15538, showed a KV allocation attempting several GiB on one specific GPU and failing there. The important diagnostic fact is the targeted device, not total VRAM across the machine.

Current multi-GPU controls include --split-mode and --tensor-split, but their behavior and maturity vary by mode and backend. Use the current llama.cpp multi-GPU documentation, not an old command copied from a discussion.

6. Read the deepest backend error

These sequences mean different things even though they end with the same llama.cpp exception.

CUDA

cudaMalloc failed: out of memory
...
failed to allocate buffer for kv cache

First suspect device headroom, context, concurrency and placement.

Vulkan

Upstream issue #9708 reported:

vk::Device::allocateMemory: ErrorOutOfDeviceMemory
...
failed to allocate buffer for kv cache

That points to the Vulkan device allocator, not a broken GGUF.

ROCm

A ROCm log may identify the exact device that rejected the allocation. That matters in asymmetric or already-loaded multi-GPU layouts.

CPU

A CPU backend can fail because the requested cache exceeds available system memory or a parameter fit path chose an unrealistic context.

OpenVINO

OpenVINO has had backend-specific allocation constraints that are not explained by ordinary aggregate RAM or VRAM math. RAMGPT already has a focused case study: llama.cpp OpenVINO Context Failure: The 4 GiB KV Cache Limit.

That page is the special-case investigation. This page is the general hub.

7. If it should fit, test for a backend limit or regression

Once you have established that:

then “out of memory” is no longer automatically the whole diagnosis.

Useful isolation tests include:

same model + same context, different backend
same backend, smaller context
same backend, --parallel 1
same configuration, older known-good llama.cpp commit
same configuration, current upstream commit
same model with less GPU offload

A clean old-vs-new commit boundary can turn a vague memory complaint into a reproducible regression.

Do not hide a backend bug by permanently shrinking everything if the allocation should fit and a known-good runtime proves that it used to.

A practical diagnostic command set

Record the binary first:

command -v llama-cli
readlink -f "$(command -v llama-cli)"
llama-cli --version

Record devices and current memory:

llama-cli --list-devices
nvidia-smi

Then test a small context:

llama-cli   -m model.gguf   --ctx-size 4096   -n 1   -p "test"

For server workloads, remove concurrency as a variable:

llama-server   -m model.gguf   --ctx-size 8192   --parallel 1

If those work, increase one dimension at a time.

Always check your binary’s current --help; llama.cpp flags and defaults evolve.

How to verify the fix

A real verification should show more than “server started.”

Confirm that the new run:

  1. uses the intended runtime/version;
  2. allocates the KV buffer successfully;
  3. reports the expected context and parallel settings;
  4. identifies the expected backend/device placement;
  5. runs a prompt through generation;
  6. survives repeated requests if the failure originally appeared only under concurrency.

If you changed multiple knobs at once, you do not know which one fixed the problem. Reintroduce settings one at a time if you need the actual root cause.

Similar-looking errors that need different fixes

Error family First diagnosis
failed to allocate buffer for kv cache KV/backend allocation and memory placement
failed to allocate compute buffers compute graph/workspace allocation, not necessarily KV
unknown model architecture runtime does not recognize the model family
tensor ... has wrong shape converter/model/runtime structural mismatch
memory grows until OOM after many tokens/requests leak, cache growth or runtime regression
backend-specific size/limit error before KV failure backend constraint may be the real root cause

For the broad classification tree, see llama.cpp Errors and Fixes. For the exact GLM5Next loader failure, see unknown model architecture ‘glm5next’.

What to include in a bug report

Do not report only “KV cache failed.”

Include:

llama.cpp version/commit:
binary path:
OS:
CPU:
GPU(s):
GPU driver:
backend:
build options:
model repository/revision:
exact GGUF filename(s):
context size:
parallel setting:
cache type K:
cache type V:
GPU layers:
split mode / tensor split:
free memory immediately before launch:
exact command:
backend allocator error immediately before KV failure:
full relevant log:
known-good older commit, if any:

The single most useful line is often the allocator failure immediately before failed to allocate buffer for kv cache. It tells you whether you are solving a capacity problem, a placement problem, or a backend-specific failure.

Sources and further reading

Continue reading