vLLM 0.30.0 FA4 SplitKV Bug: H100 Decode Can Be 48% Slower
A Sep 30 vLLM report traces a large H100 decode regression to FA4 ignoring SplitKV num_splits on SM90; the one-line kernel fix is merged upstream but not yet pinned by vLLM.
Approximately 11 min read
A performance regression that returns the right tokens is easy to miss.
That is what makes vLLM issue #59362 interesting.
On September 30, 2026, an H100 user reported that vLLM 0.30.0 could make Gemma 4 decode almost 48% slower than vLLM 0.26.0 at roughly 5.9k tokens of context, even though time to first token stayed flat and output correctness was not the problem.
The failure was narrowed to FlashAttention 4 on SM90.
The kernel accepts a SplitKV value called num_splits, but the SM90 path effectively runs as if the split count were always one. The extra split CTAs still launch, but most of them receive empty work ranges.
So the GPU looks busy enough from the outside.
The output remains correct.
The decode step is simply much slower than it should be.
The particularly awkward part is release timing: the one-line fix was merged into vllm-project/flash-attention on September 28 as PR #198, but at publication time vLLM main still pins commit 9cd61de, one commit before the fix.
This is a source analysis of the upstream issue, the FlashAttention patch, and the vLLM dependency pin. RAMGPT did not independently run the H100 benchmark.
The symptom is decode slowdown, not a crash
The reporter tested google/gemma-4-26B-A4B-it on one H100 80GB with the FlashAttention backend.
At short context, the regression is visible but modest.
At longer context, it gets much larger.
The upstream end-to-end measurements are median milliseconds per decode step:
| Context | Batch | vLLM 0.26.0 | vLLM 0.30.0 | 0.30.0 + fix |
|---|---|---|---|---|
| ~260 tokens | 1 | 4.335 ms | 4.691 ms | 4.349 ms |
| ~260 tokens | 4 | 5.179 ms | 5.558 ms | 5.382 ms |
| ~5.9k tokens | 1 | 4.378 ms | 6.460 ms | 4.440 ms |
| ~5.9k tokens | 4 | 5.462 ms | 7.455 ms | 5.568 ms |
Those are measurements reported in vLLM issue #59362.
RAMGPT did not run them.
At the ~5.9k batch-1 point, the issue describes the v0.30.0 decode step as about 47.5% slower than v0.26.0.
After mounting the one-line fix over the JIT FA4 kernel, the reported decode step moves back to within roughly 1.4% of the old version for that case.
That is unusually strong localization for a performance regression.
Time to first token stays flat
The reporter first separated prefill from decode.
Time to first token was about 61 ms on both versions.
Per-token decode was the part that moved.
That matters because an end-to-end request benchmark can hide the distinction:
prompt processing
+ scheduling
+ model decode
+ output handling
= total request time
If only decode regresses, changing tokenizer settings, prompt preprocessing, or scheduler knobs is unlikely to fix the underlying cause.
The issue then tested batch size, model runner version, text-only inputs, MoE settings, context length, attention backend, kernel profiles, and finally the FA4 SplitKV parameter itself.
The search space collapses quickly once those controls are separated.
The context-length dependence points directly at attention
The most revealing A/B is context length.
The reported v0.26.0 decode step stays nearly flat:
~260 tokens: 4.335 ms
~5.9k tokens: 4.378 ms
The reported v0.30.0 step does not:
~260 tokens: 4.691 ms
~5.9k tokens: 6.460 ms
That shape suggests work scaling with the KV sequence rather than a fixed overhead in the scheduler or model runner.
The issue author then switches v0.30.0 to TRITON_ATTN and reports that roughly 70% of the gap disappears.
That is not a recommendation that Triton is globally better.
It is a very useful diagnostic because it removes FA4 while keeping the rest of the vLLM stack much closer to the same configuration.
The kernel profile isolates one operation
At batch 1 around 5.9k tokens, the reporter says the attention portion moves from about:
0.53 ms per decode step
to
2.51 ms per decode step
while GEMM, MoE, and normalization work remain essentially flat.
The global-layer FA4 kernel is even more specific.
The reported launch time changes from roughly:
18.5 microseconds
to
311 microseconds
with the same grid, thread count, and shared-memory configuration.
This is the kind of performance bug that can survive a superficial kernel-launch comparison.
The same number of CTAs can launch while the useful work is distributed very differently.
What SplitKV is trying to do
During single-token decode, attention often has a lot of KV history and very little query work.
For a simplified grouped-query attention shape, imagine:
query tokens: 1
KV heads: 2
KV sequence: thousands of tokens
GPU: 132 SMs
If one CTA handles the entire KV range for each KV head, there may be very little parallel work available.
SplitKV divides the KV sequence into multiple ranges:
KV range
-> split 0
-> split 1
-> split 2
-> ...
-> split N
Each split computes a partial attention result.
A combine step then merges those partial results correctly.
The reason to split is parallelism.
Long KV sequences can be distributed across more SMs instead of making a tiny number of blocks do all of the work serially.
Gemma 4 exposes the failure clearly
The issue uses Gemma 4 26B-A4B.
Its global-attention layers have only two KV heads per sequence in the tested shape.
The reporter points out that without useful KV splitting, only two of the H100’s 132 SMs have meaningful work for that part of the decode.
That makes the missing parallelism unusually visible.
This does not mean the bug is a Gemma 4 bug.
The bug is in the SM90 FA4 SplitKV work distribution.
Gemma 4 is simply a particularly effective workload for exposing it because the combination of few KV heads and a long KV sequence makes the lost split parallelism expensive.
The actual bug is one missing constructor argument
The source-level mechanism is surprisingly small.
The affected FA4 code uses a BlockInfo object to determine which part of the KV range each split should process.
The method still receives num_splits as an argument.
But on the non-packed SM90 path, the method reads the object’s stored split count instead.
Conceptually:
caller passes num_splits = 16
method receives num_splits = 16
but BlockInfo.num_splits was left at default 1
method uses BlockInfo.num_splits
effective split count = 1
The SM100 path already constructs BlockInfo with the runtime split count.
The SM90 path did not.
PR #198 adds the missing value to the SM90 constructor.
The relevant change is effectively:
BlockInfo(
...
num_splits=num_splits,
)
One argument restores the intended work distribution.
Why the result is still correct
A performance failure without a correctness failure can look suspicious.
Why do the extra splits not corrupt attention?
Because the unused splits end up with empty KV ranges.
The combine kernel knows how to handle them.
Conceptually:
split 0 -> real KV range
split 1 -> empty
split 2 -> empty
...
combine -> empty splits contribute zero weight
So the final attention result can still be correct while most of the supposed parallel work is wasted.
This is an important systems lesson:
Correct numerical output does not prove the optimized execution path is actually doing the intended work.
Performance-sensitive code needs performance regression tests as well as correctness tests.
The kernel-only sweep makes the bug obvious
The issue includes a small kernel reproduction that holds the attention shape constant and changes num_splits.
For a head dimension of 512, two KV heads, batch 1, and roughly 7k KV tokens, the reported timings are:
| num_splits | v0.26.0 | v0.30.0 | v0.30.0 + fix |
|---|---|---|---|
| 1 | 385.4 us | 360.1 us | 358.6 us |
| 4 | 104.7 us | 360.2 us | 101.1 us |
| 16 | 33.0 us | 356.5 us | 32.7 us |
| 64 | 20.3 us | 363.0 us | 20.2 us |
Again, these are upstream measurements.
The pattern is more important than the absolute numbers.
In v0.26.0:
more splits
-> more useful parallelism
-> lower latency
In unpatched v0.30.0:
more splits
-> almost no change
After the patch:
the scaling returns
That is a clean signature of an ignored runtime tuning parameter.
Why passing num_splits manually does not help
If the SM90 BlockInfo object still stores one, changing the external num_splits value does not solve the bug.
The issue explicitly lists manual num_splits as something that does not work.
That can save a lot of pointless tuning.
If your build contains the affected FA4 code, trying:
num_splits = 4
num_splits = 16
num_splits = 64
may leave the measured kernel time essentially flat.
The problem is not choosing a better split count.
The problem is that SM90 is not honoring the count in the place that controls work distribution.
The regression arrived through a dependency pin
This is the operationally interesting part.
vLLM did not need a large attention rewrite in its own repository to inherit the problem.
PR #54819 advanced the pinned vllm-project/flash-attention revision.
The issue’s release map is:
vLLM 0.26.0 through 0.29.0
-> older FlashAttention revisions
-> this specific bug not present
vLLM 0.30.0
-> pin 506341a
-> bug present
vLLM main at publication time
-> pin 9cd61de
-> still missing fix 6d11b3c
The fix itself was already merged into the FlashAttention fork on September 28.
The serving runtime had simply not advanced its pin far enough yet.
That gap is a classic source of release asymmetry:
upstream subproject fixed
!=
downstream package already fixed
Fixed upstream is not the same as fixed in the vLLM wheel or container you are running.
Check the exact pin, not just the vLLM version string
If you are investigating this regression, inspect both layers.
First:
python -c "import vllm; print(vllm.__version__)"
Then inspect the source revision or build metadata used by the package or container.
For a source checkout, check:
git rev-parse HEAD
grep -n "GIT_TAG" cmake/external_projects/vllm_flash_attn.cmake
At publication time, vLLM main still points that GIT_TAG at:
9cd61de38763d712bb6ce56e2a02cc2bf718c89f
The merged FlashAttention fix is:
6d11b3c1b6ca7dac0b111bb0db63ca16eeb19a46
Do not assume those values will remain current after this article is published.
The important method is to check the actual dependency pin in the build you are deploying.
A practical diagnostic sequence
If H100 decode becomes slower after moving to vLLM 0.30.0, especially with long context and few KV heads, use a narrow A/B.
1. Separate TTFT from decode
Measure prefill and per-token decode independently.
If TTFT is stable while decode grows, keep attention and KV-dependent paths high on the suspect list.
2. Sweep context length
Use the same model and batch size at a short and longer context.
A gap that grows with KV length is consistent with lost decode-attention parallelism.
3. Compare attention backends
As a diagnostic, try TRITON_ATTN on the same vLLM release.
The upstream reporter recovered a large part of the gap that way.
Do not treat that as a universal permanent recommendation; use it to determine whether FA4 is the differentiating component.
4. Check the FA4 pin
Confirm whether your build contains FlashAttention PR #198 or a later revision that includes it.
5. Re-run the same workload after the fix
Keep model, batch, context, quantization, GPU, and scheduling settings unchanged.
Otherwise a fixed result can simply be a different benchmark.
Do not confuse this with every H100 decode regression
The issue itself points to another vLLM performance regression involving the V2 model runner on Confidential Computing VMs.
That is a different problem.
For this FA4 SplitKV case, the reporter specifically controlled for the runner:
V2 runner -> slow
V1 runner -> similarly slow
The issue also rules out multimodal inputs, structured output, MoE tuning files, and ordinary batching as the primary cause for the tested setup.
The high-value signature is narrower:
H100 / SM90
+ FA4
+ vLLM 0.30.0-era FlashAttention pin
+ long decode context
+ low KV-head parallelism
+ num_splits sweep has no effect
That is much more useful than saying vLLM 0.30 is slow.
The broader lesson is dependency-level performance correctness
Inference runtimes are stacks.
A simplified vLLM path can include:
vLLM scheduler
-> model runner
-> attention backend
-> pinned kernel project
-> CuTe / CUTLASS
-> CUDA
-> GPU
A one-line constructor omission several layers down can erase a major optimization while leaving every output numerically correct.
That is why production inference regression testing should include at least four dimensions:
correctness
latency
throughput
memory
A release can pass the first one and still be unacceptable.
For this specific case, the evidence is unusually crisp:
On the upstream reporter’s H100 workload, vLLM 0.30.0’s SM90 FA4 path effectively collapses SplitKV to one useful split because BlockInfo does not receive the runtime num_splits value. FlashAttention PR #198 adds that missing argument, restores split scaling in the contributor’s kernel tests, and returns Gemma 4 decode close to the older vLLM baseline.
The remaining deployment question is not whether the one-line fix exists.
It is whether the exact vLLM build you run has actually picked it up.