llama.cpp 'failed to allocate compute pp buffers': How to Fix It
Fix llama.cpp compute pp buffer allocation failures by shrinking micro-batches, preserving device-memory headroom, checking backend limits, fit warnings, and speculative decoding.
Approximately 7 min read
If llama.cpp ends with graph_reserve: failed to allocate compute buffers followed by failed to allocate compute pp buffers, do not assume the GGUF is broken and do not treat it as the same problem as a KV-cache allocation failure.
Use this order:
- Read the backend allocator line immediately above the final error.
- Reduce physical micro-batch size with
-ub/--ubatch-size. - Reduce logical batch size if necessary.
- Leave more device-memory headroom by changing model placement/offload.
- Check whether
--fitactually succeeded instead of merely being enabled. - If speculative decoding or MTP is active, test without it.
- On multi-GPU systems, inspect the device receiving the failed allocation rather than aggregate VRAM.
- If an older known-good commit works with the same configuration, treat the change as a possible regression.
The final line names the subsystem that failed. The more useful cause is often one or two lines earlier.
What the compute pp buffer is telling you
llama.cpp needs device memory for more than weights and KV cache:
model weights
+ KV cache
+ output buffers
+ prompt-processing compute graph
+ backend workspaces
+ speculative/draft context allocations
+ driver/runtime overhead
A prompt-processing graph can require a large temporary compute arena. The model and KV cache can load successfully, then context initialization can fail when the backend cannot reserve that arena.
That is why “the model fits” is not enough.
Read the deepest allocator error first
Upstream issue #14836 is a clear Vulkan example. The run used two 48 GB Radeon Pro W7900 GPUs. The log showed model loading and a roughly 752 MiB KV cache, but the failing request was a single Vulkan buffer of about 4.76 GB. The backend reported that the requested buffer exceeded its device-memory allocation limit before llama.cpp emitted the generic compute-buffer error.
The lesson is not that every Vulkan setup has that limit. It is that the backend-specific line can distinguish a capacity problem from a per-allocation limit.
Preserve messages such as:
CUDA cudaMalloc failed: out of memory
Vulkan requested buffer exceeds device allocation limit
ROCm allocation failed on device N
fit logic says projected usage exceeds free memory
Do not report only the final failed to allocate compute pp buffers line.
Reduce physical micro-batch first
Current llama.cpp CLI documentation distinguishes:
-b, --batch-size logical maximum batch size
-ub, --ubatch-size physical maximum batch size
The current documented defaults are 2048 logical and 512 physical. The physical micro-batch is a direct compute-graph sizing lever.
A conservative diagnostic run is:
llama-cli -m model.gguf -c 8192 -b 512 -ub 128 -n 1 -p "test"
If the failing production command used -ub 4096 or -ub 8192, lower it before assuming the GPU or GGUF is defective.
Watch the allocation size printed by the backend. If lowering -ub makes the requested compute arena smaller and the run succeeds, you have isolated a graph-sizing problem.
Reduce logical batch if needed
If lowering -ub is not enough, lower -b too. The goal is not to prescribe a universal optimum. It is to find a small configuration that initializes, then increase one dimension at a time until the failure boundary is visible.
This makes the bug report much more useful than “it OOMs with my normal settings.”
Leave VRAM headroom
A model can fit in VRAM while leaving too little room for the compute arena.
Current llama.cpp exposes -ngl / --gpu-layers / --n-gpu-layers for model placement. Reducing GPU-resident weights can be a useful diagnostic because it leaves more device memory for runtime allocations.
This is a trade:
more weights in VRAM -> less headroom for compute buffers
fewer weights in VRAM -> more headroom, but potentially slower execution
Optimize the whole workload, not just “maximum GPU layers.”
Check whether –fit actually fitted the configuration
Current llama.cpp enables --fit by default and documents it as adjusting unset arguments to fit device memory. It also exposes a target margin and a minimum context.
Read the fit log instead of assuming the feature succeeded.
Upstream issue #27872 shows a recent failure mode. With --cpu-moe, the run first warned that projected device usage exceeded free VRAM, then logged that fitting could not proceed because user tensor-buffer overrides were already set. The server continued with the oversized settings and later failed at compute-buffer allocation.
So:
--fit enabled
!=
fit completed successfully
Look earlier in the log for projected memory, target headroom, and warnings saying the fitting pass aborted.
Speculative decoding can add another compute arena
MTP and other speculative paths can create additional runtime state.
In upstream issue #27282, a Qwen3.8-27B run on an RTX 4090 loaded the target context, then failed while creating the native MTP draft context because the draft side requested another CUDA compute allocation. The report also isolated micro-batch sizing as one factor in the allocation size.
For diagnosis, compare the same model and context with speculation disabled and enabled. Current llama.cpp exposes speculative modes through --spec-type, including draft-mtp.
If the target model runs normally and the failure appears only when a draft context is enabled, focus on draft compute memory, its device placement, and current upstream behavior.
Also test a real completion. Server startup or a health endpoint does not necessarily prove that every graph needed for generation has already been allocated.
Multi-GPU total VRAM can mislead you
A single compute buffer must be allocated on a specific backend device according to the graph and split strategy. Free memory on another GPU does not automatically satisfy that request.
Record per-device state:
llama-cli --list-devices
nvidia-smi
Then match the allocator failure to the device named in the log.
Do not sum all GPU memory and treat it as one flat pool.
Compute buffers are not KV cache
The two error families can appear close together but point to different first levers:
failed to allocate buffer for kv cache
-> context, KV datatype, sequences, KV placement
failed to allocate compute pp buffers
-> prompt graph, micro-batch, compute arena, backend allocation
There is overlap because both consume device memory, but if the KV cache already allocated and the next failure is a large graph buffer, -ub is a more direct first test than changing KV precision.
For the other branch, use llama.cpp ‘failed to allocate buffer for kv cache’: How to Fix It.
Test for a regression when the configuration should fit
Once you have a controlled reproduction, compare a known-good build with a current build using the same GGUF, context, batch, micro-batch, backend, offload, and speculative settings.
If the older commit succeeds and the newer one requests a much larger compute arena or fails at a new point, preserve the first known-good and first known-bad commits. That turns a vague OOM complaint into actionable regression evidence.
Practical isolation sequence
Start small:
llama-cli -m model.gguf -c 4096 -b 512 -ub 128 -n 16 -p "test"
Then increase one dimension at a time: context, micro-batch, logical batch, GPU offload, speculation, and production concurrency. Stop at the first transition that reproduces the failure and preserve the allocator line.
What to include in a bug report
Record the llama.cpp version/commit, binary path, OS, backend, GPU and driver, model revision, exact GGUF, context, batch, ubatch, GPU layers, split settings, --fit output, speculation mode, free device memory, exact failed allocation size, target device, and any known-good/known-bad commits.
How to verify the fix
Do not stop at “the server stayed up.” Verify that model weights load, KV cache allocates, compute pp buffers allocate, a real prompt completes prompt processing, token generation begins, and the production speculation/concurrency mode survives if that was where the failure occurred.
The useful rule is: treat failed to allocate compute pp buffers as a graph-allocation problem until the log proves otherwise. Read the backend failure, then reduce the parameters that directly control the compute arena before changing unrelated model settings.