vLLM DeepSeek-V4 Startup Crash: Stale max_model_len After KV Auto-Fit
A fresh vLLM fix traces a DeepSeek-V4 CUDA-graph startup assertion to one stale max_model_len copy left behind after KV-cache auto-fitting reduces context.
Approximately 13 min read
A startup crash that appears only when a GPU has slightly less free memory is easy to blame on CUDA graphs, a flaky allocator, or a giant context window.
A new vLLM fix opened on September 22 points to a more precise failure: the engine reduces max_model_len to fit the available KV cache, but one cached copy of the old length survives inside ModelState.
That stale integer is enough to make two parts of the same runtime disagree about the largest sequence they are preparing for.
On the reported DeepSeek-V4 sparse-MLA setup, that disagreement reaches a hard assertion during CUDA-graph capture:
max_model_len: reduced from 1048576 to 980224 to fit in available GPU memory
...
File ".../vllm/models/deepseek_v4/sparse_mla.py", in _build_c128a_metadata
assert active_topk_width >= cm.max_seq_len // self.compress_ratio
AssertionError
The proposed patch in vLLM PR #58149 is tiny: when GPUModelRunner.update_max_model_len() receives the auto-fitted length, update self.model_state.max_model_len too.
The interesting part is not the line count. It is the lifecycle bug behind it:
load model
-> cache original context length in several objects
-> profile KV memory
-> reduce context length
-> update some cached copies
-> forget one copy
-> construct attention metadata from the new value
-> CUDA-graph capture reads the old value
-> dimensions disagree
-> startup assertion
This article is a source-level analysis of vLLM issue #57034 and PR #58149. RAMGPT did not reproduce the DeepSeek-V4 run. Hardware behavior, memory figures, and the original reproduction below are from the upstream issue reporter. The PR’s regression test is a CPU-only state-propagation test.
The reported configuration
Issue #57034 uses a development vLLM build whose relevant paths were reported to match upstream main at the time of investigation.
The environment was:
hardware: 2 x NVIDIA GB10
compute capability: 12.1
PyTorch: 2.13.0+cu130
CUDA: 13.0
Python: 3.12.3
model: DeepSeek-V4-Flash-0731
tensor parallel: 2
speculative decoding: DSpark, k=5
max model length: auto
full model context: 1,048,576
KV cache dtype: fp8
gpu memory utilization: 0.85
The important setting is not speculative decoding. It is:
--max-model-len auto
When vLLM profiles available memory and discovers that the KV cache cannot support the checkpoint’s full context, it is allowed to reduce the effective maximum length.
In the failing startup, the log reports:
max_model_len: reduced from 1048576 to 980224
The reporter measured about 9.81 GiB of available KV cache in that run.
On another startup with about 12.3 GiB available, no reduction was needed and the same configuration started successfully.
That is why the bug can look nondeterministic.
The model and command do not have to change. A small difference in free GPU memory can decide whether auto-fit executes, which decides whether the stale-state path exists.
Auto-fit is supposed to make startup safer
The purpose of context auto-fitting is straightforward.
Suppose the model advertises:
max context = 1,048,576 tokens
but the KV cache that fits in the current GPU memory budget can support only:
980,224 tokens
A reasonable runtime can choose:
requested/advertised maximum
|
v
profile available KV capacity
|
v
safe effective maximum = 980,224
That should turn a hard memory failure into a smaller but usable context window.
The bug in #57034 is not that vLLM computes the reduced value incorrectly.
The reduced value is propagated correctly to some runtime state. The problem is that not every consumer switches to the same new value.
One number existed in several places
Modern inference engines are full of derived state. A configuration value such as max_model_len does not remain a single field read from one global object forever.
Different components cache it because they need fast access or because they are constructed at different phases of startup.
The issue traces three important copies.
First, the V2 GPU runner has its own maximum length:
GPUModelRunner.max_model_len
Second, request state tracks a maximum:
req_states.max_model_len
Third, the model-state object caches the value when it is constructed:
ModelState.max_model_len
Before PR #58149, the runner’s update method handled the first two:
self.max_model_len = max_model_len
self.req_states.max_model_len = max_model_len
but not the third.
So after auto-fit, the state could become:
model_config.max_model_len = 980224
GPUModelRunner.max_model_len = 980224
req_states.max_model_len = 980224
ModelState.max_model_len = 1048576 <- stale
That is the whole bug in one table.
Why ModelState has the old value
The order of construction matters.
ModelState is created during model loading, before KV-memory profiling has finished deciding whether the full context fits.
At construction, it copies the then-current value from the model config:
1,048,576
Later, vLLM profiles memory and reduces the model configuration to:
980,224
The worker broadcasts that fitted value and calls the runner’s update_max_model_len() method.
But because ModelState already copied the original integer into its own field, changing the config object later does not retroactively change the cached copy.
This is a classic snapshot problem:
A = config.max_model_len
later:
config.max_model_len = smaller_value
A is still the old integer
Nothing about Python references saves you here because the cached field is a scalar value, not a live property pointing back to the config.
Why DeepSeek-V4 sparse MLA exposes the mismatch
A stale field is only dangerous if two later computations use different copies.
DeepSeek-V4’s sparse-MLA metadata path provides exactly that situation.
The attention metadata builders are constructed after auto-fitting. They therefore see the reduced model configuration and size their C128A structures for 980,224 tokens.
But CUDA-graph capture later asks ModelState for the worst-case sequence length. That field still says 1,048,576.
So one side prepares capacity for:
980224
while the capture path asks it to represent:
1048576
The assertion is therefore not random. It is a consistency check catching two different views of the same supposed runtime limit.
The numbers line up exactly
The issue provides enough information to follow the failed inequality.
For the reduced context, the metadata builder computes its compressed capacity around a compression ratio of 128.
The reported calculation is:
c128a_max_compressed
= align_up(cdiv(980224, 128), 128)
= 7680
During capture, the stale ModelState provides:
max_seq_len = 1048576
and:
1048576 / 128 = 8192
The active width is clamped by the capacity built from the reduced value:
active_topk_width <= 7680
but the assertion requires enough width for the stale capture length:
active_topk_width >= 8192
Those conditions cannot both be true.
So the runtime reaches:
7680 >= 8192
and fails.
That exact numerical contradiction is useful because it distinguishes this problem from a vague “DeepSeek CUDA graph issue.” The assertion is doing its job: two subsystems were configured with incompatible maximum lengths.
Why the crash happens during CUDA-graph capture
CUDA graphs try to capture reusable execution shapes so later work can launch with less CPU overhead.
For capture, the runtime often constructs worst-case or representative metadata rather than processing an ordinary live request.
The model-state capture path needs a maximum sequence length to describe that worst case.
Before the patch, it reads:
ModelState.max_model_len
which is precisely the stale field.
That explains the timing:
model loads
KV cache profiles
context auto-fits successfully
attention metadata builds
warmup/capture begins
assertion fires
engine dies before serving traffic
The failure is therefore a startup crash, not a decode-time out-of-memory condition.
Why more free memory can make the bug disappear
This is one of the best diagnostic clues.
If there is enough KV capacity for the full advertised context, auto-fit does nothing:
original = 1048576
fitted = 1048576
Now even a stale copy is harmless because every copy has the same value.
If memory pressure forces a reduction:
original = 1048576
fitted = 980224
then the stale copy becomes observable.
So a pattern like this is plausible:
run A: slightly more free VRAM -> starts
run B: slightly less free VRAM -> max_model_len reduced -> assertion
That can look like GPU instability if the operator does not notice the auto-fit log line.
The first thing to compare between successful and failed startups should therefore be the resolved context length, not only the final traceback.
The proposed fix is deliberately small
PR #58149 changes GPUModelRunner.update_max_model_len() to fan the fitted value into ModelState too:
self.max_model_len = max_model_len
self.req_states.max_model_len = max_model_len
self.model_state.max_model_len = max_model_len
The principle is more important than the syntax:
if a runtime API promises to update a derived configuration value, every cached copy that affects execution must participate in that update contract.
The PR author notes that the older V1 runner already updates another derived value, effective_drafter_max_model_len, through this method. So propagating to model_state follows an existing pattern rather than inventing a separate mechanism.
The regression test checks fan-out, not GPU execution
PR #58149 adds a CPU-only test around the update method.
It creates a stand-in runner where all cached lengths begin at:
FULL_LEN = 1048576
then calls:
update_max_model_len(980224)
The first assertion checks that model_state.max_model_len changed.
The second test checks all represented cached copies and fails if any still carries the original number.
The PR reports:
without source change: 2 failed
with source change: 2 passed
These are upstream PR test results, not RAMGPT measurements.
Also note what the test does not prove. It does not load DeepSeek-V4, allocate a KV cache, or reproduce the CUDA-graph assertion on a GPU. It verifies the state-propagation mechanism that the issue identified as the cause.
That is appropriate for a regression test on a simple setter, but it should not be misreported as an end-to-end hardware reproduction.
The PR author checked current main’s source path
The original issue came from a development build and explicitly said the reporter had not independently reproduced the bug on a fresh upstream-main binary.
PR #58149 closes part of that uncertainty by inspecting current upstream source at commit 362a64b21d and confirming the three required conditions still existed before the patch:
ModelState caches max_model_len
GPUModelRunner.update_max_model_len does not update ModelState
capture reads ModelState.max_model_len
That is strong source evidence that the stale-state path exists on main.
It is still different from saying the PR author performed the original two-GB10 end-to-end reproduction on main. The article should keep those claims separate.
This can affect more than one model-state capture class
DeepSeek-V4 sparse MLA produces the visible assertion, but PR #58149 points out that the stale max_model_len field is consumed by multiple model-state capture paths:
default.py
encoder_decoder.py
encoder_only.py
mamba_hybrid.py
That does not prove all of those paths currently crash.
It means the stale value is part of shared state used beyond one DeepSeek-specific class. DeepSeek-V4’s C128A assertion is the loud symptom because its buffer capacity calculation exposes the mismatch directly.
This is another useful debugging pattern: the first model that trips an invariant is not necessarily the only model touched by the underlying state bug.
Why rope_state was not changed
The PR also mentions another object initialized with the pre-auto-fit context: rope_state.
Why not update everything blindly?
Because reducing the maximum context leaves that particular structure over-allocated, not under-sized. An over-sized RoPE state may waste memory, but it does not create the same reduced-buffer-versus-larger-request contradiction in this path.
So the patch stays narrow rather than broadening into unrelated lifecycle changes.
That restraint matters in runtime fixes. Once a stale-state bug is found, it is tempting to rewrite configuration ownership everywhere. A minimal repair plus a test for the exact update contract is easier to review and less likely to create another regression.
The immediate workaround does not require patching vLLM
Issue #57034 reports a practical workaround:
set an explicit --max-model-len that already fits the KV cache
The reporter used:
--max-model-len 900000
The exact safe value is hardware- and configuration-dependent. 900000 is not a universal recommendation.
The point is to avoid this state transition:
auto starts at huge value
-> model state caches huge value
-> KV profiling shrinks config later
If the runtime begins with an explicit length that fits, the objects are constructed with one consistent maximum and no auto-fit fan-out is required.
For operators, that is a useful mitigation while an upstream fix is still open.
A diagnostic checklist for this exact failure
If vLLM dies during warmup or CUDA-graph capture with a DeepSeek-V4 sparse-MLA assertion, check the startup log in this order:
1. Was --max-model-len set to auto or -1?
2. Did vLLM print "max_model_len: reduced from ... to ..."?
3. Does the traceback end in sparse_mla.py::_build_c128a_metadata?
4. Is the failure an AssertionError rather than a direct CUDA OOM?
5. Does an explicit smaller --max-model-len make startup reliable?
6. Which vLLM commit contains or lacks PR #58149?
If the successful startup shows no reduction but the failed startup does, that correlation is particularly strong.
Do not reduce the diagnosis to “the context is too large.” Auto-fit already recognized that and chose a smaller context. The bug is that one consumer failed to adopt the chosen value.
This differs from the recent speculative-draft context crash
RAMGPT’s September 21 Technical Daily covered a different vLLM length mismatch: a target model accepted positions beyond a speculative draft model’s shorter RoPE cache, eventually surfacing as a CUDA illegal-memory-access error.
This new bug has a different contract and a different failure stage:
Sep 21 draft mismatch:
target context > draft positional capacity
-> long live request
-> draft path overruns its supported range
Sep 22 auto-fit mismatch:
old cached max_model_len > newly fitted max_model_len
-> startup CUDA-graph capture
-> metadata capacities disagree
-> AssertionError
Both involve “context length,” but combining them into one article would hide the actual search intent. One is a two-model speculative-decoding compatibility problem; the other is stale configuration state inside one runtime after memory profiling.
The deeper lesson: mutable config creates a synchronization contract
A configuration value feels static until the runtime changes it after inspecting hardware.
The moment vLLM auto-fits max_model_len, the field becomes dynamic state.
That creates a contract:
profile hardware
-> choose fitted value
-> publish fitted value
-> every execution consumer must observe it
Caching is safe only if one of two things is true:
A. the source value can never change after construction
or:
B. every cache participates in a reliable invalidation/update mechanism
The failed path had neither property. ModelState cached a value before profiling, and the later update did not reach it.
This class of bug is not unique to context length. Similar failures can appear anywhere a serving engine snapshots values that are later auto-tuned:
batch limits
block counts
workspace sizes
parallelism-derived shapes
cache capacities
capture dimensions
The safest tests are often not enormous GPU integration tests. A small state-consistency test can protect the invariant directly.
Current status
As of September 22, 2026, vLLM PR #58149 is open and unmerged.
The PR is mergeable at the time of this analysis and includes a CPU regression test. That does not mean a released package already contains it.
For an affected deployment, verify the exact vLLM commit rather than assuming latest includes an open pull request.
Until the fix is merged into the revision you deploy, an explicit context length that is known to fit the KV-cache budget is the upstream reporter’s documented workaround.
Practical takeaway
If a DeepSeek-V4 vLLM server sometimes starts and sometimes dies in _build_c128a_metadata during CUDA-graph capture, look for the line that says the context was auto-reduced.
The reported failure is not simply:
not enough VRAM
It is:
not enough KV capacity for full context
-> vLLM correctly reduces max_model_len
-> one ModelState copy stays at the original length
-> metadata is sized for the smaller value
-> capture requests the larger value
-> assertion
PR #58149 fixes that specific state-propagation hole by updating model_state.max_model_len together with the runner and request-state copies.
That one assignment restores the invariant the rest of the startup pipeline assumed all along: there should be one effective maximum context length, even after hardware-driven auto-fitting changes it.