llama.cpp MTP Regression: CUDA MoE Fusion Breaks Exactness
A Sep 20 llama.cpp report bisects an MTP speculative-decoding regression to CUDA MoE weighted-reduction fusion: draft acceptance drops from 0.82 to 0.48 and MTP becomes slower than no draft.
Approximately 10 min read
A speculative-decoding optimization is only useful if the target model still defines the answer.
That is why a fresh llama.cpp report from September 20, 2026 is more serious than an ordinary performance regression.
The reporter bisected a Gemma 4 MoE + MTP setup to one CUDA change:
b10750 = good
b10751 = bad
The single commit between those builds is 3466812, merged through PR #25952:
cuda: fuse MoE weighted expert reduction
On the reported RTX 2070 SUPER setup, the change did three things at once:
draft acceptance: 0.823 -> 0.481
MTP throughput: 46.35 -> 35-37 tok/s
greedy equivalence: identical -> different output
The surprising part is that ordinary non-speculative decoding stayed roughly unchanged.
So this is not a broad “CUDA got slower” regression.
It is a much narrower interaction between:
MoE target model
+ CUDA weighted-reduction fusion
+ MTP speculative decoding
This article is a source analysis of issue #29168 and the code introduced by PR #25952. The benchmark numbers below are upstream measurements, not RAMGPT results.
The exact upstream configuration
The report uses:
GPU: RTX 2070 SUPER, sm_75, 8 GB
OS: Windows
CUDA package: 12.4 prebuilt llama.cpp binaries
Target: Gemma-4-26B-A4B-it QAT GGUF
Draft: Gemma-4 MTP sidecar GGUF
Partial offload: -ngl 31 --n-cpu-moe 21
The speculative path is enabled with an MTP draft:
llama-server \
-m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf \
--model-draft MTP/mtp-gemma-4-26B-A4B-it-Q4_0.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
-ngl 31 \
--n-cpu-moe 21 \
-c 32768 \
-ctk q8_0 \
-ctv q8_0 \
-fa on \
-b 512 \
-ub 512
The reporter compares that against the same target model with the draft flags removed.
Sampling is greedy:
{
"n_predict": 128,
"temperature": 0,
"top_k": 1,
"seed": 1,
"cache_prompt": false
}
That makes the correctness test unusually clean.
If speculative decoding is exact, the drafted and non-drafted paths should produce the same final greedy output for the same target model.
What changed across one build
The issue gives this comparison:
| Metric | b10750, last good | b10751 and later |
|---|---|---|
| Draft acceptance | 0.823 | 0.481 |
| Mean accepted length | 2.65 | 1.95 |
| Throughput with MTP | 46.35 tok/s | 35.36-36.73 tok/s |
| Throughput without MTP | 38.10 tok/s | 37.02-38.33 tok/s |
| Drafted output equals greedy output | yes | no |
The key control is the non-draft row.
Without speculative decoding, target-model throughput is almost flat across the boundary:
38.10 tok/s before
37.02-38.33 tok/s after
That sharply narrows the search space.
If this were a general CUDA matmul regression, a weight-loading problem, or a broad Gemma 4 slowdown, the ordinary target path should also move substantially.
It does not.
MTP changes from an acceleration into overhead
Before the regression:
no draft: 38.10 tok/s
MTP: 46.35 tok/s
The draft is doing useful work because enough proposed tokens survive target verification.
After the regression, the relationship reverses:
no draft: 37.02-38.33 tok/s
MTP: 35.36-36.73 tok/s
Now the draft model consumes compute, but the target rejects enough proposals that the speculative machinery does not recover its cost.
This is the operational consequence of acceptance rate.
A draft model is not automatically a speedup.
A rough mental model is:
speculative benefit
≈ accepted target tokens per verification step
minus draft/verification overhead
If acceptance collapses, speculative decoding can become slower than ordinary decoding even if both the draft and target run correctly in isolation.
Why output equality matters more than tokens per second
The speed regression is useful because it makes the bug visible.
But the stronger signal is this:
b10750:
drafted greedy output == non-drafted greedy output
b10751+:
drafted greedy output != non-drafted greedy output
Speculative decoding is supposed to change how efficiently the target result is produced, not redefine the target distribution.
With greedy decoding, that expectation is especially straightforward.
Conceptually:
target model says token X
A draft can propose candidates faster, but verification should still reduce the process to the target model’s decision.
If enabling the draft changes the final greedy text, performance is no longer the only concern.
That is a correctness failure in the observed configuration.
The bisect points to PR #25952
The issue author tested eight builds during the bisect:
b10723 GOOD
b10736 GOOD
b10740 GOOD
b10749 GOOD
b10750 GOOD
b10751 BAD
b10775 BAD
b10837 BAD
The acceptance signal was sharply separated:
good: ~0.8229
bad: ~0.4806
There is exactly one commit between b10750 and b10751:
3466812
That commit is the merge of PR #25952.
The report also says the behavior is still present in b10964, b11056, and b11057.
That makes this a useful search asymmetry: someone debugging low MTP acceptance today might naturally inspect the draft model, quantization, prompt, or --spec-draft-n-max, while the reported first bad change is in the target model’s CUDA MoE reduction path.
What PR #25952 actually optimized
Before the fusion, the MoE combine tail conceptually did something like:
expert outputs
-> multiply each expert by router weight
-> write weighted results
-> reduce expert contributions
PR #25952 replaced that sequence with a fused CUDA kernel that performs weighting and reduction together.
The PR describes its main goal as removing intermediate global-memory traffic and reducing two physical kernels to one.
The merged implementation contains a kernel shaped like:
float sum = first_expert * first_scale * first_weight;
for (int expert = 1; expert < n_expert_used; ++expert) {
sum += expert_value * scale * weight;
}
dst[...] = sum;
That is attractive for performance.
The original PR reported prompt-processing improvements of roughly 3.6% to 7.1% across several tested model/hardware combinations, along with passing backend tests and unchanged WikiText-2 perplexity in its validation setup.
Those are upstream PR measurements.
The important point is that the original testing did not reveal the new MTP exactness case later reported on a Turing GPU with partial MoE offload.
The original PR said numerics were unchanged
This is where the new report becomes interesting.
PR #25952 explicitly reported:
WikiText-2 PPL identical
llama-cli output unchanged
no changes to weights or numerics
That was true for the configurations the PR author tested.
Issue #29168 provides a different configuration where drafted greedy output is no longer identical across the speculative and non-speculative paths.
Both observations can coexist.
A numerical optimization can pass ordinary output/PPL tests on one set of models and hardware while exposing a sensitivity in a different path.
The missing test dimension appears to be something closer to:
MoE + CUDA fusion + speculative target verification
rather than ordinary single-path generation.
The floating-point explanation is plausible, but not proven
The issue author suspects the fused reduction changes floating-point accumulation behavior enough to perturb target logits.
That is plausible because floating-point arithmetic is not perfectly associative:
(a + b) + c
is not guaranteed to produce exactly the same rounded result as:
a + (b + c)
And a fused kernel can change more than source-level order. It can change where intermediate values live, when they are rounded, and whether compiler instructions combine operations.
A tiny logit change normally does not matter.
But speculative verification can be sensitive around token boundaries.
If two candidate logits are extremely close, a tiny numerical difference can change the argmax:
before:
token A = 12.500001
token B = 12.500000
some changed arithmetic:
token A = 12.499999
token B = 12.500000
Now the target chooses B instead of A.
That can cause a draft proposal to be rejected earlier, which changes acceptance statistics and can alter the greedy generated sequence.
However, the upstream issue labels this as a suspected mechanism. No maintainer fix or confirmed root-cause patch was attached when this article was written.
So the correct status is:
bisected regression: strong evidence
floating-point mechanism: plausible hypothesis
final upstream diagnosis: pending
Why PPL can miss this class of problem
Perplexity is an aggregate metric.
Speculative exactness is a path-equivalence property.
Those are different tests.
A kernel can produce tiny numerical differences that barely move aggregate perplexity while still flipping a small number of greedy token decisions.
That means a validation matrix such as:
PPL unchanged
backend ops pass
ordinary CLI output looks fine
is valuable but not sufficient to prove:
speculative decode output == target-only output
For speculative systems, that equivalence deserves a dedicated regression test.
A stronger regression test is simple
The issue’s test pattern is worth keeping.
For a deterministic prompt:
1. Run target-only greedy decoding.
2. Run the same target with MTP enabled.
3. Hash the final output bytes.
4. Compare acceptance statistics.
For example:
curl -s localhost:8080/completion \
-d '{"prompt":"Write one vivid paragraph about the sea.","n_predict":128,"temperature":0,"top_k":1,"seed":1,"cache_prompt":false}' \
| jq -r .content \
| sha256sum
The important assertion is not:
output looks reasonable
It is:
speculative output hash == target-only output hash
That catches exactly the property speculative decoding is supposed to preserve.
What to check if your MTP acceptance suddenly collapses
If MTP used to help and now becomes slower, separate four questions.
1. Does target-only decoding still behave normally?
Measure without the draft.
If target-only throughput and output are stable while MTP falls apart, the problem is probably not generic model execution.
2. Compare greedy output with and without the draft
Use deterministic settings:
temperature = 0
top_k = 1
fixed prompt
fixed build
Then compare exact output, not just semantic similarity.
3. Read draft acceptance
The issue’s boundary is dramatic:
~0.82 -> ~0.48
A sudden acceptance cliff can be more diagnostic than overall tok/s.
4. Record the exact llama.cpp build
For this report:
last good: b10750
first bad: b10751
first bad commit: 3466812
That is far more useful than saying “a recent version is slower.”
Current diagnostic switches and workarounds
The safest operational workaround from the issue is simple:
disable MTP speculative decoding
On the affected reporter’s builds, target-only decoding is faster than the broken MTP path and avoids the speculative output-divergence problem.
For debugging the suspected fusion interaction, current llama.cpp CUDA source also contains a broad environment switch:
GGML_CUDA_DISABLE_FUSION=1
That disables CUDA fusion generally, not just this one MoE reduction. It is therefore useful as an A/B diagnostic, not automatically the best production configuration.
A clean test matrix would be:
A. target only
B. target + MTP
C. target + MTP + CUDA fusion disabled
If B diverges from A but C returns to A, that strengthens the fusion diagnosis on your machine.
Do not treat that as a confirmed upstream fix. As of September 20, issue #29168 is open and no dedicated repair PR was linked from it.
This is not the same as RAMGPT’s earlier scheduler-hash crash coverage
RAMGPT recently covered a separate speculative-decoding failure where scheduler capacity could trigger an assertion/crash.
This new issue has a different search intent:
scheduler-capacity bug:
process fails visibly
#29168:
process runs, but acceptance collapses and greedy equivalence breaks
That distinction matters for debugging.
A crash pushes you toward bounds, scheduler state, and hard failures.
A low-acceptance exactness regression pushes you toward numerical path differences between target-only and speculative execution.
The broader runtime lesson
Kernel fusion is normally discussed as a throughput optimization:
fewer launches
less memory traffic
more locality
But inference runtimes increasingly contain algorithms that depend on relationships between multiple model evaluations.
Speculative decoding is one of them.
That changes the standard for an optimization.
It is not enough to ask:
Is the fused kernel numerically close?
You may also need to ask:
Does it preserve the target decisions required by the higher-level algorithm?
For ordinary generation, a tiny floating-point change may be invisible.
For speculative verification, the same tiny difference can turn a proposed token from accepted to rejected and erase the entire speedup.
Practical takeaway
If llama.cpp MTP on a CUDA MoE model suddenly shows both:
much lower draft acceptance
and
MTP slower than target-only decoding
check the build boundary and compare exact greedy output with and without speculation.
For issue #29168, the upstream reporter found:
b10750: acceptance 0.823, MTP faster, output identical
b10751+: acceptance 0.481, MTP slower, output diverges
and bisected the boundary to the merged CUDA MoE weighted-reduction fusion in commit 3466812 / PR #25952.
That bisect is strong evidence.
The proposed floating-point explanation is not yet a confirmed root cause, so the responsible conclusion is narrower:
on the reported Gemma 4 MTP + RTX 2070 SUPER configuration, enabling the post-b10750 CUDA path breaks the exact target-only behavior that speculative decoding is expected to preserve.
Until upstream resolves it, target-only decoding is the conservative workaround, and GGML_CUDA_DISABLE_FUSION=1 is a useful diagnostic lever for isolating the fusion path.