Troubleshooting

llama.cpp RX 7800 XT: Windows Vulkan MTP Falls 54%

A Sep 25 llama.cpp report finds MTP throughput on an RX 7800 XT dropping from 55.79 to 25.80 t/s across a 146-commit range; Linux RADV does not reproduce it.

Approximately 11 min read

A performance patch can be excellent on the hardware it was tuned for and still expose a nasty portability edge somewhere else.

That is the interesting part of llama.cpp issue #29410, opened September 25, 2026.

The issue reporter compared two llama.cpp revisions on an AMD Radeon RX 7800 XT under Windows. Prompt processing barely moved, ordinary token generation slowed, and MTP speculative decoding at a 48k-token depth fell from a reported 55.79 tokens per second to 25.80 tokens per second.

The obvious suspect is tempting: the newer build includes the freshly merged Vulkan int8 cooperative-matrix work from PR #27952, which was designed specifically to accelerate quantized matrix multiplication on AMD RDNA3 and RDNA4.

But the issue is not a bisect.

The two builds are 146 commits apart.

And a second tester on an RX 7900 XTX under Linux RADV could not reproduce the regression.

So the useful conclusion today is narrower than “PR #27952 broke AMD Vulkan.”

It is this:

A new Windows RX 7800 XT report shows that the recent RDNA3 Vulkan performance path is not producing the same gains across driver, GPU, and workload combinations, and the MTP slowdown is large enough to justify a proper matrix of A/B tests before blaming the model or speculative decoder.

That is a much better debugging starting point.

The exact reported setup

Issue #29410 compares these revisions:

older:
c77ae695c
2026-09-17 era

newer:
84e76d8a2
2026-09-24 era

The reporter explicitly notes that the newer revision is 146 commits after the older one. That matters because it prevents a causal shortcut: every performance change in that range remains a candidate until narrower testing eliminates it.

The machine is:

OS: Windows 11
GPU: AMD Radeon RX 7800 XT 16 GB
architecture: RDNA3
Vulkan driver: AMD proprietary
CPU: Ryzen 5 7600X
RAM: 32 GB
reported shared memory: 32768 bytes
matrix cores: KHR_coopmat

The model is described in the issue as a Qwen3.5/3.6 35B-A3B MoE model in Q4_K_M, with native MTP available.

For the empty-context benchmark, the reporter uses a llama-bench command equivalent to:

device: Vulkan0
GPU layers: 99
CPU MoE layers: 18
Flash Attention: on
prompt: 512 tokens
generation: 32 tokens

For the long-context MTP measurements, the server is configured with:

context: 163840
KV cache: q4_0 K + q4_0 V
MTP draft depth: 3
reasoning: off
batch / ubatch: 2048 / 2048

The long-context measurements are taken after filling the KV cache to approximately 48k and 130k tokens.

That is important context. This is not a tiny synthetic matrix-kernel result. It is a serving-style path with MoE, quantized KV, MTP, and substantial context depth.

What the reporter measured

These are the upstream issue author’s numbers. RAMGPT did not run this benchmark.

Test c77ae695c 84e76d8a2
pp512 1174 t/s 1151 ± 9 t/s
tg32, no MTP 42.4 t/s 36.1 ± 0.6 t/s
MTP n=3, tg128 at 48k 55.79 t/s 25.80 t/s
MTP n=3, tg128 at 130k 39.15 t/s 27.14 t/s

The 48k MTP point is the headline-sized change:

55.79 -> 25.80 t/s

That is a reduction of about 54%.

But the shape of the whole table is more informative than the largest percentage.

Prompt processing is roughly flat:

1174 -> 1151 t/s

Ordinary decode is slower:

42.4 -> 36.1 t/s

MTP decode is much slower at both measured depths.

That means this is not simply “the new fast prefill kernel failed to improve pp512.”

There is a decode-side regression to explain too.

Acceptance stayed at 1.00

This detail sharply separates #29410 from the CUDA MTP exactness regression RAMGPT covered on September 20.

In that earlier case, draft acceptance collapsed and greedy output diverged. A speculative path that proposes bad candidates naturally loses performance because the target rejects more of the draft.

Issue #29410 says something different:

MTP acceptance = 1.00
old build: 1.00
new build: 1.00

The draft tokens are still being accepted.

So the giant drop at 48k is not explained by “MTP suddenly became a bad drafter.”

Conceptually, the old run looks like:

draft useful tokens
-> target verifies them
-> accepted
-> speedup

The new run still has:

draft useful tokens
-> target verifies them
-> accepted

but each speculative step is taking much longer.

That moves the investigation away from model quality and toward execution cost:

kernel selection
tile shape
driver behavior
MoE dispatch
memory traffic
long-context attention cost
work scheduling

This distinction is exactly why acceptance rate belongs beside tokens per second in speculative-decoding diagnostics.

PR #27952 is relevant, but it is not yet convicted

The newer revision contains merged PR #27952:

vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4

The design makes sense.

Before the change, llama.cpp’s Vulkan cooperative-matrix path mainly relied on fp16 matrix operations after dequantizing weights into shared memory.

Many AMD GPUs have strong int8 matrix throughput, so the new path performs quantized MMQ work with int8 cooperative-matrix instructions.

The PR supports several common quant formats, including:

Q4_0
Q4_1
Q5_0
Q5_1
Q8_0
Q3_K
Q4_K
Q5_K
Q6_K
MXFP4
NVFP4
IQ4_NL

That makes it directly relevant to the Q4_K_M model in #29410.

The upstream PR benchmarks also show why this work was attractive.

On a Radeon 8060S Strix Halo system, the PR author reported Qwen3.6-35B-A3B Q4_K_M pp512 moving from:

777.8 -> 1023.6 t/s

about a 1.32x improvement.

On an RDNA4 Radeon AI PRO R9700, the same model/quant row was reported as:

2760 -> 3365 t/s

about a 1.22x improvement.

Those are PR-author measurements, not RAMGPT measurements, and they are not the same hardware or driver as issue #29410.

That difference is the entire point.

The Windows result does not match the Strix Halo story

The RX 7800 XT report expected some benefit because both systems sit in the broader RDNA3 family and the new path is explicitly enabled there.

Instead:

Strix Halo PR benchmark:
Q4_K_M MoE pp512 -> strong gain

RX 7800 XT Windows issue:
Q4_K_M MoE pp512 -> essentially flat
decode -> slower
MTP -> much slower

This is a classic case where “same architecture family” is not enough information.

The actual execution environment also includes:

integrated vs discrete GPU
shared-memory limits
driver implementation
wave/tile choices
MoE workload shape
context depth
quant layout
kernel dispatch thresholds

A shader that is well-tuned for one RDNA3 configuration can choose a less favorable tile or occupancy point on another.

That does not prove that this is what happened here. It tells us what kind of hypothesis is worth testing.

Linux RADV provides an important negative control

A second commenter tested another RDNA3-family card:

GPU: Radeon RX 7900 XTX
OS: Linux
driver: Mesa RADV
model: dense Qwen3.8-27B IQ3_S

Their interleaved llama-bench runs compared a pre-#27952 build with a current master containing the new path.

They reported:

Test Before #27952-era merge Master containing #27952
pp512 943.9 t/s 980.2 t/s
tg128 51.11 t/s 50.98 t/s

For their MTP probe, median generation throughput was reported as:

79.6 -> 82.0 t/s

with overlapping acceptance ranges.

In other words: no comparable decode or MTP regression.

That does not exonerate every commit in the 146-commit range, because the cross-check also changes the GPU, operating system, driver, model type, quant, and workload.

But it is useful evidence against a universal statement like:

"RDNA3 Vulkan is now slower"

The failure surface is narrower.

The 32 KB shared-memory clue

The most specific upstream hypothesis so far comes from the contributor behind #27952.

Responding to the Windows report, they noted that the optimization was tuned and tested on Linux, and called out the Windows device’s reported shared-memory limit:

Windows RX 7800 XT:
shared memory = 32768

The Linux RADV cross-check reports:

shared memory = 65536

The contributor suggested that the 32 KB limit may restrict the Windows path to a medium tile and that it might need a smaller L-warptile.

That is a useful source-level hypothesis because tile size, cooperative-matrix packing, and shared-memory budget are directly connected.

But it is still a hypothesis.

There is no upstream A/B in #29410 yet showing:

same RX 7800 XT
same Windows driver
same build
only tile selection changed
-> regression disappears

Until that test exists, the responsible wording is:

The shared-memory/tile explanation is plausible and comes from the patch author, but it is not yet a confirmed root cause.

Why the 146-commit range matters so much

The issue author is careful about this, and we should be too.

Between the old and new revisions, several Vulkan changes landed. The report itself mentions work involving IQ formats, symbol visibility, Intel Xe Flash Attention, Adreno, convolution alignment, and other MMQ changes.

Some appear unlikely to be hot for this exact Q4_K_M workload.

“Unlikely” is not a bisect.

A clean performance investigation should reduce the range rather than reason from commit titles alone.

At minimum, I would want these four builds:

A. last known-good revision
B. immediately before #27952 merge
C. #27952 merge commit
D. current master

Then run the exact same command and warmup procedure.

That separates:

regression introduced before #27952
regression introduced by #27952
regression introduced after #27952
interaction among multiple changes

Without that split, reverting one attractive suspect and seeing a speed change can still mislead if surrounding changes interact.

A better A/B matrix for AMD Vulkan

If I had access to the affected machine, I would keep the test small and systematic.

1. Separate prefill from decode

Run llama-bench with:

pp512
tg32 or tg128

Do not summarize them into one “overall speed” number.

The issue already shows that prefill and decode can move in different directions.

2. Test MTP with acceptance beside throughput

Record:

effective tg
draft acceptance
mean accepted length
context depth

If acceptance remains at 1.00 while throughput changes, the problem is execution cost rather than draft quality.

3. Hold context depth fixed

The issue reports both 48k and 130k.

That is useful because long context changes attention and KV behavior. A regression that appears only at one depth may point somewhere very different from a regression that is flat across all depths.

4. Compare dense and MoE models

The Linux negative control used a dense model while the Windows failure uses a 35B-A3B MoE model.

A same-card dense-vs-MoE comparison would help answer whether the slow path is tied to expert matmuls or is more general.

5. Compare the exact #27952 boundary

This is the most valuable missing test.

Everything else is still inference until the commit range gets smaller.

Do not confuse this with the earlier Vulkan prefill fast-path bug

RAMGPT has already covered a Vulkan case where a layout guard silently disabled a fast Flash Attention path for hybrid KV-cache views.

That September 2 article was about:

valid KV-cache view
-> guard rejects it
-> fast path not selected
-> prefill much slower

Issue #29410 has a different search intent:

fresh RDNA3 int8 coopmat path
+ Windows proprietary driver
+ RX 7800 XT
+ MoE Q4_K_M
+ MTP / decode slowdown

Keeping those separate matters for both debugging and search.

“Vulkan performance regression” is too broad a category to be useful.

The hardware, driver, quant, model architecture, and phase of inference are part of the bug report.

Do not confuse it with the CUDA MTP exactness regression either

The September 20 CUDA MoE article had this signature:

acceptance collapses
greedy output changes
MTP becomes slower

Today’s Windows Vulkan report has:

acceptance stays 1.00
output-quality failure not reported
MTP becomes much slower

Those are different failure classes.

The first points toward numerical/exactness behavior in verification.

The second points toward per-step runtime cost.

That one metric — acceptance — changes the investigation.

What operators should do today

If you run llama.cpp on AMD Vulkan and a recent update makes MTP or decode slower, record the exact revision before changing model settings.

Then capture:

GPU model
OS
Vulkan driver
reported shared memory
model architecture
quant type
context depth
CPU-MoE setting
MTP depth
acceptance
pp and tg separately

If the regression resembles #29410, compare a known-good build against a build immediately around the #27952 merge instead of simply disabling MTP and concluding that speculative decoding is broken.

For production use, the conservative choice is the revision that is already measured as stable on your own workload.

That is not a recommendation to pin forever. It is a way to keep serving predictable while the upstream range is narrowed.

The broader lesson: performance patches have an environment

PR #27952 is exactly the kind of optimization inference runtimes need.

It turns available int8 matrix hardware into useful work for quantized models and, in its reported Linux benchmarks, produces substantial gains.

Issue #29410 is not evidence that the optimization was a mistake.

It is evidence that performance is a function, not a property:

performance =
    kernel
  x hardware
  x driver
  x tile choice
  x model shape
  x quant
  x runtime path

“RDNA3 support” is not one machine.

“Vulkan” is not one driver.

“MTP” is not one cost profile.

The fresh Windows RX 7800 XT report matters because it gives upstream a concrete counterexample with exact commands and measurements.

Now the next useful step is not more speculation.

It is a tighter bisect and an A/B around the shared-memory and tile-selection hypothesis.

Until that arrives, the strongest statement supported by the sources is straightforward:

on one reported Windows RX 7800 XT + Q4_K_M MoE configuration, a Sep 24 llama.cpp build shows much slower decode and MTP than a Sep 17 build, while a Linux RADV RDNA3 cross-check does not reproduce the slowdown.

Sources and further reading

Continue reading