Local AI

vLLM Qwen3.8 NVFP4: Prefix-Cache Hits Corrupt DFlash2/DSpark Output

A vLLM 0.30/0.31 report isolates output corruption to prefix-cache hits combined with hidden-state speculative drafters on Qwen3.8-27B NVFP4.

Approximately 6 min read

A new vLLM bug report is unusually useful because it does more than say “speculative decoding broke.”

The reporter isolated a specific interaction:

Qwen3.8-27B target
+ compressed-tensors NVFP4 checkpoint
+ hidden-state speculative drafter
  - DFlash2 or DSpark
+ prefix-cache hit
+ vLLM 0.30 or 0.31
=
coherent early requests, then corrupted repeated-token output

The same conversation replay reportedly remains coherent on vLLM 0.29. The same target without speculative decoding passes. The target with its own MTP head passes. An FP8 target with the same DFlash2 configuration passes.

That is a much narrower failure surface than “Qwen3.8 is unstable.”

What the failure looks like

In issue #60174, the first requests in a long multi-round agent conversation are normal.

Once later requests resume from cached prefix blocks, the model begins emitting one repeated token instead of coherent text. The reporter also sees the speculative acceptance length collapse to 1.00, meaning the drafter is contributing no accepted speculative progress.

The test uses five consecutive requests from the same conversation. Prompt lengths rise from roughly 31K to 67K tokens, with 25 tool schemas and a fixed seed. Each request is sent to a fresh engine in sequence so later requests can reuse the earlier prefix.

This matters because the failure is not a cold-start problem. It appears after cache reuse becomes active.

Why prefix caching changes the problem

Automatic prefix caching stores reusable inference state for a prompt prefix so a later request can skip recomputing that prefix.

For a normal transformer, people often think of this as reusing KV-cache blocks.

Hybrid models complicate that picture. Qwen3.8 includes linear-attention / GDN-style state in addition to ordinary attention state. vLLM therefore has to decide where a cached prefix is safe to resume and which recurrent state must be reconstructed or preserved.

The current vLLM design documentation describes block-based prefix reuse and special handling for Mamba-like state.

That means a cache hit is not merely a performance event. It is also a state-restoration boundary.

If any component restores a state that is numerically valid but semantically mismatched, the engine can continue running while generating wrong output.

That is more dangerous than a clean crash.

The strongest part of the report: the ablation table

The reporter changes one subsystem at a time.

On vLLM 0.31:

baseline DFlash2 + NVFP4 + prefix hit          FAIL
apply neighbouring Mamba cache fix            FAIL
switch GDN prefill backend                     FAIL
switch NVFP4 linear backend                    FAIL
switch to W4A16 weight-only path               FAIL
shrink KV pool / context                       FAIL
disable speculative decoding                   PASS
use checkpoint MTP head                        PASS
use DSpark hidden-state drafter                FAIL
use FP8 target with same DFlash2 config        PASS
force cache miss with unique cache_salt        PASS

The exact numbers and outcomes are upstream-reported, but the structure of the experiment is strong.

Several obvious explanations become less likely:

The interaction requires a hidden-state drafter, a real prefix-cache hit, and the affected compressed-tensors target path.

DFlash and DSpark are an important clue

DFlash is not just a small autoregressive draft model.

vLLM’s speculators documentation describes DFlash as a draft model conditioned on hidden states from the target. It predicts a block of candidate tokens using target-derived context features.

That design can be fast because the drafter gets richer information than raw token IDs.

It also means hidden-state continuity matters.

If the target resumes from a cached prefix but the drafter’s expected recurrent or hidden-state boundary does not correspond to that cached state, the system can be internally inconsistent even though each tensor has a valid shape.

The issue reporter suspects bad hybrid-model state surviving at the cached tail boundary. That is a hypothesis, not yet a confirmed root cause, but the ablations make it a reasonable place for maintainers to inspect.

Why MTP passing does not clear speculative decoding

The checkpoint’s own MTP path reportedly passes all five requests.

That should not be simplified to:

MTP good, external speculators bad.

The more precise interpretation is that the failure does not affect every speculative architecture equally.

MTP, DFlash2 and DSpark can have different state interfaces, cache assumptions and draft-model loading paths.

So the useful diagnostic question is not “is speculative decoding enabled?”

It is:

which speculative method?
which target quantization/loader?
which cache hit path?
which hybrid-state boundary?

NVFP4 is part of the interaction, but probably not the kernel bug

The reporter tried multiple linear backends and even a weight-only alternative. The failure remained.

The FP8 target, however, passed with the same DFlash2 setup.

That points toward a difference in the target checkpoint or loader path rather than a single low-level NVFP4 matrix kernel.

The issue calls out the compressed-tensors configuration, including ignored unquantized linear-attention components alongside quantized groups.

That is a reasonable area to inspect because a hidden-state drafter depends on the target’s internal representation, not only its final logits.

But the report does not prove that the compressed-tensors loader is the root cause.

A safer production diagnostic

If a long-running agent starts producing repeated or nonsensical output only after several turns, do not immediately blame sampling temperature or the prompt.

Test the interaction directly.

1. Reproduce from a fresh engine

Replay the same ordered conversation from an empty cache.

Record the exact request number where corruption begins.

2. Force cache misses

If the runtime supports a per-request cache salt, use a unique value.

If corruption disappears only when reuse is prevented, you have separated cache-resume behavior from ordinary decode.

3. Disable the drafter

Run the same target and same prompts without speculative decoding.

If the output becomes stable, continue narrowing.

4. Switch speculative method

If the checkpoint supports native MTP, compare it with DFlash/DSpark under the same requests.

A method-specific difference is far more actionable than “spec decode fails.”

5. Compare target checkpoint formats

If an FP8 or BF16 target exists, compare it with the compressed-tensors NVFP4 target.

Keep the drafter, prompt order and cache settings fixed.

What operators should log

A production incident like this needs more than the final text.

Capture:

The key is to make cache state part of the incident record.

Without that, a failure that appears on request three may be mistaken for a random model-quality problem.

The operational lesson

Prefix caching and speculative decoding are both performance optimizations, but once they share hidden or recurrent state they can become correctness-sensitive infrastructure.

A useful deployment gate for new speculative paths is therefore not just:

cold prompt -> coherent output

It should include:

multi-turn replay
+ actual prefix-cache hits
+ long prefixes
+ tool schemas if used in production
+ repeated runs across the target quantization formats you serve

A fast path that passes single-turn smoke tests can still fail at the state-resume boundary.

For agent workloads, that boundary deserves explicit regression coverage.

Sources and further reading

Continue reading