vLLM Qwen3.8 NVFP4: Prefix-Cache Hits Corrupt DFlash2/DSpark Output
A vLLM 0.30/0.31 report isolates output corruption to prefix-cache hits combined with hidden-state speculative drafters on Qwen3.8-27B NVFP4.
Approximately 6 min read
A new vLLM bug report is unusually useful because it does more than say “speculative decoding broke.”
The reporter isolated a specific interaction:
Qwen3.8-27B target
+ compressed-tensors NVFP4 checkpoint
+ hidden-state speculative drafter
- DFlash2 or DSpark
+ prefix-cache hit
+ vLLM 0.30 or 0.31
=
coherent early requests, then corrupted repeated-token output
The same conversation replay reportedly remains coherent on vLLM 0.29. The same target without speculative decoding passes. The target with its own MTP head passes. An FP8 target with the same DFlash2 configuration passes.
That is a much narrower failure surface than “Qwen3.8 is unstable.”
What the failure looks like
In issue #60174, the first requests in a long multi-round agent conversation are normal.
Once later requests resume from cached prefix blocks, the model begins emitting one repeated token instead of coherent text. The reporter also sees the speculative acceptance length collapse to 1.00, meaning the drafter is contributing no accepted speculative progress.
The test uses five consecutive requests from the same conversation. Prompt lengths rise from roughly 31K to 67K tokens, with 25 tool schemas and a fixed seed. Each request is sent to a fresh engine in sequence so later requests can reuse the earlier prefix.
This matters because the failure is not a cold-start problem. It appears after cache reuse becomes active.
Why prefix caching changes the problem
Automatic prefix caching stores reusable inference state for a prompt prefix so a later request can skip recomputing that prefix.
For a normal transformer, people often think of this as reusing KV-cache blocks.
Hybrid models complicate that picture. Qwen3.8 includes linear-attention / GDN-style state in addition to ordinary attention state. vLLM therefore has to decide where a cached prefix is safe to resume and which recurrent state must be reconstructed or preserved.
The current vLLM design documentation describes block-based prefix reuse and special handling for Mamba-like state.
That means a cache hit is not merely a performance event. It is also a state-restoration boundary.
If any component restores a state that is numerically valid but semantically mismatched, the engine can continue running while generating wrong output.
That is more dangerous than a clean crash.
The strongest part of the report: the ablation table
The reporter changes one subsystem at a time.
On vLLM 0.31:
baseline DFlash2 + NVFP4 + prefix hit FAIL
apply neighbouring Mamba cache fix FAIL
switch GDN prefill backend FAIL
switch NVFP4 linear backend FAIL
switch to W4A16 weight-only path FAIL
shrink KV pool / context FAIL
disable speculative decoding PASS
use checkpoint MTP head PASS
use DSpark hidden-state drafter FAIL
use FP8 target with same DFlash2 config PASS
force cache miss with unique cache_salt PASS
The exact numbers and outcomes are upstream-reported, but the structure of the experiment is strong.
Several obvious explanations become less likely:
- not simply one NVFP4 GEMM kernel;
- not simply the GDN prefill backend;
- not simply the KV pool being too large;
- not all speculative decoding;
- not prefix caching by itself;
- not DFlash2 alone.
The interaction requires a hidden-state drafter, a real prefix-cache hit, and the affected compressed-tensors target path.
DFlash and DSpark are an important clue
DFlash is not just a small autoregressive draft model.
vLLM’s speculators documentation describes DFlash as a draft model conditioned on hidden states from the target. It predicts a block of candidate tokens using target-derived context features.
That design can be fast because the drafter gets richer information than raw token IDs.
It also means hidden-state continuity matters.
If the target resumes from a cached prefix but the drafter’s expected recurrent or hidden-state boundary does not correspond to that cached state, the system can be internally inconsistent even though each tensor has a valid shape.
The issue reporter suspects bad hybrid-model state surviving at the cached tail boundary. That is a hypothesis, not yet a confirmed root cause, but the ablations make it a reasonable place for maintainers to inspect.
Why MTP passing does not clear speculative decoding
The checkpoint’s own MTP path reportedly passes all five requests.
That should not be simplified to:
MTP good, external speculators bad.
The more precise interpretation is that the failure does not affect every speculative architecture equally.
MTP, DFlash2 and DSpark can have different state interfaces, cache assumptions and draft-model loading paths.
So the useful diagnostic question is not “is speculative decoding enabled?”
It is:
which speculative method?
which target quantization/loader?
which cache hit path?
which hybrid-state boundary?
NVFP4 is part of the interaction, but probably not the kernel bug
The reporter tried multiple linear backends and even a weight-only alternative. The failure remained.
The FP8 target, however, passed with the same DFlash2 setup.
That points toward a difference in the target checkpoint or loader path rather than a single low-level NVFP4 matrix kernel.
The issue calls out the compressed-tensors configuration, including ignored unquantized linear-attention components alongside quantized groups.
That is a reasonable area to inspect because a hidden-state drafter depends on the target’s internal representation, not only its final logits.
But the report does not prove that the compressed-tensors loader is the root cause.
A safer production diagnostic
If a long-running agent starts producing repeated or nonsensical output only after several turns, do not immediately blame sampling temperature or the prompt.
Test the interaction directly.
1. Reproduce from a fresh engine
Replay the same ordered conversation from an empty cache.
Record the exact request number where corruption begins.
2. Force cache misses
If the runtime supports a per-request cache salt, use a unique value.
If corruption disappears only when reuse is prevented, you have separated cache-resume behavior from ordinary decode.
3. Disable the drafter
Run the same target and same prompts without speculative decoding.
If the output becomes stable, continue narrowing.
4. Switch speculative method
If the checkpoint supports native MTP, compare it with DFlash/DSpark under the same requests.
A method-specific difference is far more actionable than “spec decode fails.”
5. Compare target checkpoint formats
If an FP8 or BF16 target exists, compare it with the compressed-tensors NVFP4 target.
Keep the drafter, prompt order and cache settings fixed.
What operators should log
A production incident like this needs more than the final text.
Capture:
- vLLM version and commit;
- target checkpoint and quantization format;
- drafter model and speculative method;
- number of speculative tokens;
- prompt token count per turn;
- prefix-cache hit size;
- Mamba/GDN cache mode;
- speculative acceptance metrics;
- generation seed and sampling settings;
- whether cache reuse was forced off;
- whether a different target format changes the result.
The key is to make cache state part of the incident record.
Without that, a failure that appears on request three may be mistaken for a random model-quality problem.
The operational lesson
Prefix caching and speculative decoding are both performance optimizations, but once they share hidden or recurrent state they can become correctness-sensitive infrastructure.
A useful deployment gate for new speculative paths is therefore not just:
cold prompt -> coherent output
It should include:
multi-turn replay
+ actual prefix-cache hits
+ long prefixes
+ tool schemas if used in production
+ repeated runs across the target quantization formats you serve
A fast path that passes single-turn smoke tests can still fail at the state-resume boundary.
For agent workloads, that boundary deserves explicit regression coverage.