Local AI

vLLM Hybrid Mamba Prefix Caching Can Cost Throughput Before a Cache Hit

vLLM issue #60008 reports an 11–16% output-throughput cost from Mamba align-mode prefix caching with zero cache hits, exposing host-side launch and prefill-splitting overhead.

Approximately 4 min read

Prefix caching is often described as an almost-free optimization: if a new request shares a prompt prefix with an old request, reuse saved state and skip repeated prefill work.

A new vLLM performance report shows why that mental model becomes more complicated for hybrid Mamba models.

Issue #60008 reports that enabling prefix caching on NVIDIA Nemotron-3.5-Lightning reduced output throughput by roughly 11–16% in a workload deliberately constructed to produce essentially no useful cache hits.

The surprising result is not that a cache miss provides no speedup. It is that maintaining the machinery required to make future Mamba hits possible added measurable work to every decode step.

The test configuration

The reporter used:

The issue was opened October 5, 2026 and remains an upstream performance report. RAMGPT has not reproduced the measurements.

Why Mamba state is different

Transformer attention can associate reusable state with KV blocks.

Mamba-style layers carry recurrent state with different structural requirements.

Current vLLM documentation describes an align cache mode. When prefix caching is enabled on affected hybrid models, vLLM stores Mamba state at scheduler/block-aligned positions so a later matching request can resume from a valid checkpoint.

That creates safe cache-hit boundaries, but it also adds work when no request actually reuses a prefix.

The reported throughput cost

The issue reports median output-throughput changes around:

concurrency 1:  about -15%
concurrency 8:  about -14%
concurrency 32: about -16%
concurrency 64: about -13%

depending on the paired comparison.

Because the prompts were random and cache hits were essentially absent, the measurement isolates the cost of keeping the Mamba prefix-cache path enabled.

Decode: small launches on the serial path

The reporter found that GPU compute time remained approximately unchanged while GPU idle time per step grew by about 0.55–0.75 ms.

One source was repeated Mamba block-table construction.

For the tested model:

4 Mamba cache groups
x 7 eager torch operations
= 28 extra eager launches per step

The report also identifies Mamba align pre-process, pre-copy and post-process kernels running outside the captured CUDA graph.

Each operation is small, but autoregressive decode repeats the path for every generated token. Fractions of a millisecond become material when they sit on that serial loop.

The GPU was waiting more, not computing much more

The report does not primarily show that prefix caching causes more neural-network math.

It shows an implementation path where host-side metadata work and small eager launches lengthen the step, increasing GPU idle time.

With DP+EP synchronization, the slowest rank can influence the step time for all ranks, amplifying small host-path differences.

Prefill: alignment can split one prompt

The report also found a separate TTFT effect.

With prefix caching off, one 8192-token prompt could be processed as one prefill step.

With the tested align-mode block size, the scheduler split it near the 4176-token boundary:

4176 tokens
+
4016 tokens

The reporter measured roughly 125 ms TTFT without that split and about 229 ms in the affected configuration.

The exact numbers are hardware-specific. The mechanism is more general: block-aligned state materialization can introduce another scheduler/prefill step, so fixed per-step overhead is paid again.

This does not mean prefix caching is bad

A cache is valuable when it hits.

Production traffic with repeated system prompts, shared documents, common agent prefixes or reused long contexts can save substantial prefill work.

The question is therefore not:

Is prefix caching faster?

It is:

At my hit rate, does saved prefill work exceed the steady-state cost of maintaining cacheable state?

A no-hit workload pays the maintenance cost but collects none of the main benefit.

Hybrid models need a three-arm benchmark

For a hybrid Mamba model, benchmark:

A: prefix caching off, no reuse
B: prefix caching on, no reuse
C: prefix caching on, realistic reuse

A versus B measures the tax.

B versus C measures the benefit from actual hits.

Without B, you cannot separate cache benefit from runtime overhead.

What operators should record

At minimum:

Tokens per second alone does not identify the cause.

Possible upstream optimization directions

Issue #60008 points toward reducing duplicated per-group arithmetic, cutting eager launches on the decode critical path, capturing suitable align operations, avoiding unnecessary pre/post work and reducing avoidable prefill splits.

Those are issue-level engineering directions, not merged guarantees.

Practical takeaway

Prefix caching can be close to a free win on some transformer workloads, but hybrid Mamba introduces state-alignment work that changes the economics.

In the reported vLLM path, a low-hit workload lost throughput before receiving any cache benefit.

If you serve a hybrid Mamba model, always benchmark prefix caching with both zero-hit and realistic-hit controls. Otherwise you cannot tell whether the cache is accelerating your traffic or merely preparing for reuse that rarely occurs.

Sources and further reading

Continue reading