Local AI

llama.cpp Vulkan Can Reuse Corrupted Scratch Data After Flash Attention

Source analysis of a Sep 28 llama.cpp Vulkan bug where flash attention or large softmax can overwrite cached scratch data, causing a later matmul to reuse corrupted input.

Approximately 6 min read

A GPU backend can return wrong intermediate data without running out of memory, crashing, or producing an obvious API error.

That is the failure mode described by llama.cpp PR #29591, opened on September 28, 2026. The patch targets a Vulkan scratch-buffer bookkeeping bug: one operation can overwrite a temporary buffer, while a later matrix multiplication still believes the old converted input is cached there.

The result is not a simple performance regression. A later matmul can reuse data that no longer represents the tensor it expects.

As of this article, PR #29591 is open. The measurements below are therefore upstream contributor results, not RAMGPT benchmarks and not proof that every Vulkan workload is affected.

The optimization that creates the hazard

The Vulkan backend uses a temporary buffer called prealloc_y. A matrix multiplication can convert an input into that buffer and remember which tensor and pipeline produced it.

That memory of the conversion is useful because a later matmul using the same input can skip doing the conversion again.

Conceptually:

matmul A
-> convert input X into prealloc_y
-> remember "prealloc_y contains converted X"

later matmul B
-> sees the remembered X
-> reuses prealloc_y

That is safe only while nothing else overwrites the buffer.

PR #29591 identifies three paths that can write different data into the same storage:

Before the proposed fix, those writes did not invalidate the record saying that prealloc_y still contained the earlier matmul input.

A subsequent matmul could therefore skip conversion and consume stale or unrelated scratch data.

Why this is a correctness bug

The dangerous state is not “buffer missing.”

It is “buffer exists, but the cache metadata lies about what is inside it.”

That distinction matters because normal allocator diagnostics may look healthy. The backend has memory. The dispatch can run. The graph can finish.

The wrong assumption is at a higher level:

bookkeeping says: converted tensor X is cached
actual buffer: flash-attention or softmax scratch data
next matmul: trusts bookkeeping and reads the buffer

This is the kind of backend bug that can be harder to notice than an OOM. A hard failure stops the request. A stale scratch reuse can instead surface as numerical error in a later operation.

Which model shapes can expose it

The PR specifically calls out graph patterns where one normalized input feeds both attention and the feed-forward path, mentioning GPT-J, Falcon, and Command-R-style models as examples.

The large-softmax trigger also depends on shape. The PR identifies its wide path at more than 16,384 columns, so long-prompt or otherwise wide workloads are more relevant than a tiny smoke test.

That does not mean every long-context Vulkan run is wrong, nor that every model in those families necessarily hits the exact sequence.

The narrower condition is:

  1. a matmul converts an input into prealloc_y;
  2. flash attention or the large-softmax path overwrites that storage;
  3. another matmul uses the same logical input;
  4. stale cache metadata lets the second matmul skip reconversion.

This sequence is what the regression test in the PR is designed to force.

The proposed fix is small because the missing invariant is small

The code change adds invalidation at the write sites.

When flash attention uses the mask optimization or sparse path, the patch clears:

prealloc_y_last_pipeline_used
prealloc_y_last_tensor_used

The large-softmax path clears the same two fields after claiming prealloc_y for its partial results.

This does not redesign the allocator. It restores a basic cache invariant:

If another operation writes into the scratch buffer, the backend must stop claiming that the previous converted matmul input is still cached there.

The patch is nine added lines in ggml-vulkan.cpp, plus the associated test changes in the PR.

Small diffs can still repair serious correctness boundaries when the failure is stale state rather than a missing algorithm.

What the upstream tests report

The contributor added a backend-ops test that performs:

matmul
-> flash attention with a large mask OR wide softmax
-> second matmul using the same input

The second weight depends on the middle operation, preventing the scheduler from simply moving work around and avoiding the intended state transition.

The PR reports the following results on AMD Radeon PRO R9700 and RX 7900 XT hardware:

master + new test: 19007 / 19011 passing
with patch:        19011 / 19011 passing

The four failures on master are reported for q4_0/q8_0 variants of the new cases.

The full Vulkan CI run still has one qwen3_0_6b save/load-state failure in both configurations, which the contributor identifies as pre-existing rather than fixed by this patch.

Those are upstream results. RAMGPT has not independently reproduced them.

The reported speed effect is effectively a non-result

The same PR includes two interleaved throughput rounds.

For Flash-Next, the contributor reports approximately:

pp512:  1253 -> 1257 t/s
pp2048: 1616 -> 1632 t/s
tg128:  51.5 -> 51.6 t/s

For Qwen3-30B-A3B:

pp512:  3728 -> 3736 t/s
pp2048: 3660 -> 3659 t/s
tg128:  144.0 -> 143.8 t/s

The PR describes these differences as within noise.

That is what we should expect from the design of the fix. Clearing two cache-tracking pointers at overwrite sites is not intended to accelerate inference. It prevents an invalid reuse.

The value is correctness, not headline tokens per second.

How I would diagnose a suspected Vulkan correctness problem

If a model behaves strangely only on Vulkan, especially at longer prompt sizes, I would avoid starting with sampling parameters.

First make the comparison deterministic enough to isolate the backend:

same model artifact
same prompt
same context length
same seed / deterministic settings where practical
same llama.cpp revision
different backend or affected-path toggle

Then reduce the graph conditions.

Useful questions include:

A difference between backends does not by itself prove this exact bug. It only narrows the investigation.

PR #29591 is most relevant when the execution graph can reuse a converted matmul input after flash-attention or wide-softmax scratch activity.

Do not convert this into a blanket “Vulkan is broken” claim

The evidence is narrower than that.

The PR identifies specific scratch-buffer reuse paths and supplies a targeted regression test. It does not establish that all Vulkan models, all GPUs, or all long-context workloads produce wrong output.

Similarly, the fact that the patch is open means operators should distinguish three states:

current build without patch
build containing the proposed patch
future upstream state if/when the PR is merged

A downstream binary can also carry the patch before an upstream release, so version labels alone may not tell the whole story.

The broader backend lesson

Scratch-buffer reuse is a performance technique, but the optimization creates a state contract.

If the backend remembers that a converted tensor lives in reusable storage, every other writer to that storage becomes responsible for invalidating the record.

That is similar to any cache:

cached value + correct invalidation -> optimization
cached value + missing invalidation -> stale state

GPU code makes the failure harder to inspect because the stale state lives inside a graph of kernels and temporary buffers rather than a normal application object.

PR #29591 is a useful reminder that backend correctness often depends on bookkeeping that looks trivial in a diff. The nine-line fix matters because it restores the relationship between what the runtime thinks is cached and what the scratch buffer actually contains.

Sources and further reading

Continue reading