Troubleshooting

llama.cpp OpenVINO Context Failure: The 4 GiB KV Cache Limit

A Sep 19 llama.cpp fix traces OpenVINO large-context failures to the backend reporting SIZE_MAX instead of the GPU's single-allocation limit, causing GGML to request one oversized KV-cache buffer.

Approximately 8 min read

A GPU can have enough memory and still refuse one allocation.

That distinction is at the center of a new llama.cpp OpenVINO failure reported on September 18 and patched in draft PR #29124 on September 19, 2026.

The visible symptom looks like an ordinary large-context failure:

failed to initialize the context
Exceeded max size of memory object allocation
failed to create context
error: unable to create context

But the root cause is more specific than “not enough VRAM.”

On the affected Intel iGPU, OpenVINO allowed a maximum single memory object of roughly 4 GiB. llama.cpp’s OpenVINO backend reported that limit to GGML as effectively unlimited, so GGML tried to allocate the entire KV cache as one object.

For the upstream Llama-3.2-3B reproduction, that single request was exactly 7 GiB.

The GPU rejected it.

This article is a source analysis of llama.cpp issue #29087 and draft PR #29124. The hardware results below are upstream measurements, not RAMGPT benchmarks.

The exact failure

The issue reproduces the bug on an Intel UHD Graphics 770 using the OpenVINO backend:

export GGML_OPENVINO_DEVICE=GPU
./llama-completion \
  -m Llama-3.2-3B-Instruct-Q4_K_M.gguf \
  -c 65536 \
  --device OPENVINO0 \
  -ngl 99 \
  -n 1 \
  -p "hi" \
  --no-warmup

The important part of the log is:

[GPU] Exceeded max size of memory object allocation:
requested 7516192768 bytes,
but max alloc size supported by device is 4294959104 bytes.

Those values are easier to understand in GiB:

requested             = 7.0 GiB
single-object limit   ≈ 4.0 GiB

Then llama.cpp surfaces the higher-level failure:

common_init_: failed to create context with model ...
llama_completion: error: unable to create context

If you search only for unable to create context, this can look like a generic model-load problem. If you search only for OpenVINO out of memory, it can look like total device-memory exhaustion.

Neither description is precise enough.

The device is rejecting one memory object that is larger than its per-allocation cap.

Single-allocation limit is not the same as total VRAM

A GPU can expose several different memory constraints:

total usable device memory
free memory at this moment
maximum size of one allocation
alignment and backend-specific allocation rules

CL_DEVICE_MAX_MEM_ALLOC_SIZE is about the largest single OpenCL memory allocation the device reports it can support.

Imagine a device that can hold 8 GiB in total but permits no individual buffer larger than 4 GiB.

These two layouts are very different:

one 7 GiB buffer
-> rejected

versus:

3.5 GiB buffer
+ 3.5 GiB buffer
-> potentially valid

The bug was not that GGML lacked the ability to work with split allocations. The bug was that the OpenVINO backend told GGML there was no meaningful maximum.

The wrong capability value was SIZE_MAX

The OpenVINO backend implements a buffer-type callback that answers a simple question:

What is the largest buffer this backend can allocate?

Before PR #29124, the OpenVINO implementation returned:

return SIZE_MAX;

In practice, that means:

Treat the maximum as effectively unlimited.

That answer is wrong for the affected Intel GPU path.

GGML already asks backends for their maximum buffer size through ggml_backend_buft_get_max_size(). The allocator uses that value when creating its chunked allocation strategy.

So GGML had the mechanism it needed. OpenVINO supplied the wrong capability metadata.

Why that turns into one giant KV-cache allocation

At long context, KV-cache memory grows with the number of cached tokens.

A simplified mental model is:

KV cache bytes ≈ context length × KV bytes per token

The exact amount depends on architecture, layer count, KV head layout, cache precision, and other details. The important property is that a longer context can push the cache across a single-object threshold even when the model weights themselves loaded successfully.

The upstream issue gives these approximate failure thresholds for full device offload on the tested iGPU with a roughly 4 GiB single-allocation ceiling:

Model Layers Reported KV bytes/token Context large enough to fail
Llama-3.2-1B 16 32 KiB 131071
Llama-3.2-3B 28 112 KiB 37449
Qwen2.5-14B 48 192 KiB 21845
Qwen2.5-32B 64 256 KiB 16383

This is why the same backend bug appears at very different context lengths for different models.

The model is not necessarily “too large” in the ordinary weight-loading sense. The context creates a KV buffer large enough to cross the device’s single-object limit.

Why reducing context works but does not fix the bug

If you reduce -c, the KV cache becomes smaller. Eventually the requested object falls below the 4 GiB limit and initialization succeeds.

That gives a tempting diagnosis:

large context = too much VRAM

But the upstream reproduction shows a more exact boundary:

large context
-> KV cache object exceeds CL_DEVICE_MAX_MEM_ALLOC_SIZE
-> allocation fails

Those are not equivalent diagnoses.

If aggregate memory is available, splitting the cache into legal-sized buffers is enough to make the large context work. That is what the proposed fix enables.

The fix is capability reporting, not a new KV-cache algorithm

PR #29124 changes a small amount of OpenVINO backend code.

During GPU initialization it queries:

CL_DEVICE_MAX_MEM_ALLOC_SIZE

with clGetDeviceInfo() and stores the result as the backend’s maximum allocation size.

Then the buffer-type callback changes conceptually from:

return SIZE_MAX;

into:

return ggml_openvino_max_alloc_size();

That is the important architectural point.

The patch does not teach GGML a brand-new split-KV-cache mechanism. Instead:

OpenVINO reports the real device limit
-> GGML allocator sees a finite maximum buffer size
-> GGML splits a large allocation into chunks
-> each individual object stays below the GPU limit

A backend capability value changes allocator behavior higher in the stack.

The allocator already understands maximum buffer size

GGML’s allocator source makes the dependency visible.

When a backend buffer type is used, GGML reads both the alignment and ggml_backend_buft_get_max_size(buft), then creates the allocator with that maximum chunk size.

That is why returning SIZE_MAX matters so much. It is not merely informational metadata printed in a log. It directly changes how large an allocation GGML believes the backend can accept.

This is a recurring inference-runtime lesson: a backend interface can contain fields that look administrative but are actually part of correctness.

Upstream validation after the change

The PR author reports retesting all four model/context combinations from the issue on the same Intel iGPU path.

These are upstream results:

Model Earlier failure threshold Context tested after patch Result
Llama-3.2-1B 131071 131072 pass
Llama-3.2-3B 37449 65536 pass
Qwen2.5-14B 21845 22016 pass
Qwen2.5-32B 16383 16384 pass

The author also tested an Intel Arc Pro B70, where the reported single-allocation limit is much larger, around 30 GiB. On that path, the tested KV caches stayed below the limit, so the allocator still used one buffer and behavior remained effectively unchanged.

CPU execution also stays on the previous unlimited-style path because this GPU-specific OpenCL query is not used there.

The OpenVINO error suggests a setting that does not solve this path

The underlying OpenVINO exception suggests enabling:

ov::intel_gpu::hint::enable_large_allocations

The issue author reports that this does not repair the failing llama.cpp KV-cache path. According to the issue analysis, the relevant flag is applied later during OpenVINO program construction, after the cache allocation that already failed.

This is a useful reminder that a lower-level library’s generic error advice may not match the calling application’s execution order.

There is a workaround, but it changes the path

The upstream issue gives this workaround:

export GGML_OPENVINO_STATEFUL_EXECUTION=1

The issue explains that this avoids moving the KV cache to the device in the same way, so the oversized GPU allocation is avoided.

That can help confirm the diagnosis, but it is not equivalent to repairing the default backend behavior. The proposed patch lets the ordinary GPU path respect the device’s actual allocation limit.

How to recognize this failure quickly

If llama.cpp with OpenVINO fails only when context becomes large, separate four questions.

1. Read the deepest exception, not only the final error

Look for:

Exceeded max size of memory object allocation

That is much more specific than:

unable to create context

2. Compare requested bytes with the device’s maximum allocation

If the log says something like:

requested: 7 GiB
max one allocation: 4 GiB

you are looking at an object-size ceiling, not automatically total-memory exhaustion.

3. See whether the boundary moves with context

Because KV memory grows with context, a clean context-size threshold is consistent with a KV allocation ceiling.

4. Check the exact llama.cpp revision

The proposed upstream fix is PR #29124.

As of September 19, 2026, the PR is still draft and unmerged. A normal git pull of upstream master therefore does not yet guarantee that the fix is present.

Verify the commit or release you are actually running.

Why this is a useful backend-debugging pattern

The visible stack is:

llama.cpp
-> GGML allocator
-> OpenVINO backend
-> OpenCL device capability
-> Intel GPU memory-object limit

The failure surfaces at the top as:

unable to create context

but the repair happens near the bottom by reporting one hardware capability correctly.

The model file is valid. The context request can be valid. The GPU may have enough aggregate memory. GGML already knows how to split buffers.

The missing piece is simply:

OpenVINO must tell GGML the real maximum size of one GPU allocation.

Practical takeaway

If OpenVINO fails at large context with:

Exceeded max size of memory object allocation
...
error: unable to create context

do not reduce the diagnosis to “VRAM OOM.”

Check the requested allocation against CL_DEVICE_MAX_MEM_ALLOC_SIZE, then check whether your llama.cpp build actually contains the OpenVINO max-allocation fix.

The proposed upstream correction is small because the allocator architecture was already capable of doing the right thing. It just needed the backend to stop saying the device’s maximum buffer size was infinite.

Sources and further reading

Continue reading