llama.cpp EAGLE-3 Bug: Stale GGML Allocation Plans Corrupt Drafts
llama.cpp issue #29313 and PR #29370 trace an EAGLE-3 acceptance collapse to stale GGML allocation plans that ignore changing tensor input/output lifetimes.
Approximately 8 min read
A graph can have the same nodes, the same tensor sizes, and still require a different memory plan.
That is the surprisingly sharp lesson behind llama.cpp issue #29313 and the new fix proposed in PR #29370.
The visible symptom is ugly for speculative decoding: EAGLE-3 draft acceptance can collapse when backend sampling is enabled. But the interesting part is lower in the stack. The failure is not primarily an EAGLE algorithm bug, a CUDA kernel bug, or a quantization problem. It is an allocator-lifetime bug.
The allocator sees a graph that looks structurally reusable. What changed is which tensors must remain alive.
That distinction is enough to corrupt draft candidates.
The short version
GGML’s graph allocator can reserve a memory layout and reuse it on later graphs when the graph shape still appears compatible.
That is normally exactly what you want. Replanning every graph would add overhead.
The problem reported in #29313 is that the reuse check did not account for two flags that change tensor lifetime:
GGML_TENSOR_FLAG_INPUT
GGML_TENSOR_FLAG_OUTPUT
An ordinary intermediate can release its storage after its last consumer.
An output cannot. Its value must remain valid after graph execution so the caller can read it.
Likewise, an input has a different lifetime because it is populated before graph execution begins.
If a tensor changes from ordinary intermediate to output while the allocator reuses a plan created under the old lifetime, two tensors that are supposed to coexist can wind up sharing storage.
Then the later write wins.
That is the bug class.
Why “same graph” is not enough
Suppose a tiny graph has three tensors:
x0 -> x1 -> x2
If x1 is only an intermediate, an allocator may legally reuse its memory once its last consumer has finished.
Now consider the same operations and the same tensor sizes, but mark x1 as an output:
x0 -> x1 -> x2
^
output: must stay alive
The topology did not change.
The tensor’s required lifetime did.
A reusable allocation plan therefore needs to encode more than dimensions and buffer capacity. It also needs to remain valid for the lifetime rules under which it was planned.
PR #29370 makes that explicit in ggml-alloc.c. The proposed patch stores the relevant input/output flags in each planned tensor allocation and compares them during the reuse check.
Conceptually, the reuse condition changes from something like:
same shape + enough space
to:
same shape + enough space + compatible lifetime flags
That last term is what was missing.
build_sampling() creates the dangerous transition
The upstream issue identifies llm_graph_context::build_sampling() as a concrete place where the lifetime set can change without the graph looking different enough to force a new allocation plan.
Only sampler results belonging to active sequences were being marked as outputs.
With multiple slots, the active set can change from one sampling step to another:
step A: slots 0, 1, 2 active
step B: slots 0, 2 active
step C: slots 1, 3, 4 active
The sampling graph can remain structurally similar while a different subset of tensors becomes externally visible output.
The stale plan still says some of those buffers may be reused early.
But the current graph now says they must survive until extraction.
Issue #29313 describes the consequence directly: output tensors that must remain live simultaneously can receive the same address, and later computation overwrites earlier output values.
For ordinary intermediate math, reuse is optimization.
For simultaneously live outputs, it is corruption.
Why EAGLE-3 makes the symptom obvious
The report uses Qwen3-1.7B Q4_K_M with a Qwen3-1.7B EAGLE-3 draft model and multiple parallel slots.
EAGLE-style speculative decoding depends on a draft path producing useful candidate tokens that the target model can accept.
If sampled candidate data is corrupted, the server does not necessarily crash. It can simply produce terrible drafts.
That is a nastier failure mode than an obvious allocator assertion because it can look like “speculative decoding is bad on this model.”
The original issue reports a 24-question MT-Bench first-turn run at temperature 0 with five requests in flight. On that reporter’s RTX 4080 SUPER setup, mean accepted length / acceptance was reported as:
backend sampling enabled: 1.33, alpha = 0.251
backend sampling disabled: 2.10, alpha = 0.53
Those are upstream reporter measurements, not RAMGPT benchmarks.
PR #29370 includes a second upstream run on an RTX 4070 Laptop with 24 prompts, -np 5, and temperature 0. Its reported draft acceptance was:
| Build | Backend sampling | Draft acceptance |
|---|---|---|
| master | on | 0.024 |
| master | off | 0.115 |
| PR #29370 | on | 0.112 |
Again, these are measurements supplied in the open PR. RAMGPT has not independently reproduced them.
What matters more than the exact numbers is the A/B shape: disabling the affected backend-sampling path largely restores acceptance, and the proposed allocator fix brings the backend-sampling result back close to that control.
The proposed fix has two layers
PR #29370 does not rely on only one tactical workaround.
It changes both the allocator and the sampling graph.
1. Make the allocator lifetime-aware
The patch adds a small helper that extracts only the flags that affect the allocation plan:
INPUT | OUTPUT
Those flags are stored when the plan is built.
When ggml_gallocr_alloc_graph() considers reusing the old plan, a changed lifetime flag invalidates reuse and forces replanning.
This is the general correctness fix. It protects callers beyond build_sampling() because the allocator itself now knows that a lifetime change matters.
The PR also adds two allocator regression tests:
ordinary tensor -> output -> ordinary tensor
ordinary tensor -> input -> ordinary tensor
The tests are designed to catch exactly the illegal overlap created when a plan is reused across changing lifetimes.
2. Keep sampling outputs stable across active-slot changes
There is still a performance concern.
If active samplers constantly change which tensors are outputs, the now-correct allocator could end up replanning frequently.
So the PR also changes build_sampling() to mark the result tensors of inactive samplers as outputs too:
sampled
probs
logits
candidates
That does not make inactive samplers execute. It makes the output-lifetime declaration stable across changes in the active sequence set.
The allocator fix protects correctness.
The sampling change reduces needless replanning by making the hot-path graph’s output contract less dynamic.
That pairing is a good design pattern: fix the generic invariant at the lower layer, then remove avoidable churn at the caller.
Why this is different from “just turn backend sampling off”
The upstream issue gives --no-spec-draft-backend-sampling as a useful control because it avoids the affected path and restores acceptance in the reporter’s setup.
That is valuable for diagnosis.
It is not the same thing as fixing the allocator.
A stale lifetime plan is a general memory-correctness problem. If another caller changes input/output status after reserve, it can create the same category of invalid reuse even if speculative decoding is not involved.
That is why PR #29370 puts the main guard in ggml-alloc.c rather than only adding an EAGLE-specific condition.
For users, disabling backend sampling can be a temporary A/B test.
For the runtime, lifetime-aware plan validity is the stronger fix.
This can masquerade as a model-quality problem
I like this bug because it demonstrates how misleading modern local-LLM symptoms can be.
Imagine the debugging sequence without the allocator evidence:
EAGLE acceptance suddenly poor
-> suspect draft model compatibility
-> suspect tokenizer mismatch
-> suspect quantization
-> suspect sampler settings
-> suspect CUDA numerical drift
All of those are plausible.
But here the candidate tensor itself can be overwritten because its storage lifetime is wrong.
The model can be fine.
The tokenizer can be fine.
The draft algorithm can be fine.
The memory plan can still make the observed acceptance look awful.
When a feature switch causes a large quality or acceptance change without an obvious model-level reason, it is worth asking whether the switch also changes graph outputs, buffer ownership, scheduling, or lifetime.
A practical diagnostic path
If I were checking a suspicious speculative-decoding regression around this issue, I would start with four controlled comparisons:
1. same model, backend sampling on vs off
2. same concurrency, especially np > 1
3. same prompt set and deterministic temperature
4. affected revision vs a build containing the allocator fix
Then I would watch two classes of signal separately:
correctness signal:
- candidate IDs
- accepted length
- acceptance rate
- deterministic output agreement
runtime signal:
- allocator replans
- buffer overlap assertions/tests
- unexpected graph reserve events
The important thing is not to conclude “backend sampling is slower” or “EAGLE is inaccurate” from one symptom. The upstream evidence points to corrupted sampling outputs caused by an invalid allocation-plan reuse decision.
The regression test is more valuable than the benchmark
The acceptance numbers make the bug visible, but the strongest part of the proposed fix is the small allocator-level test.
The test does not require a real model, a GPU benchmark, or EAGLE-3.
It asks a simpler invariant:
If a tensor becomes an output, can a later tensor still overlap it?
The correct answer is no.
That is a much better long-term guard than relying only on a full server benchmark whose acceptance can move for many unrelated reasons.
PR #29370 reports that the new input/output flag tests fail on master and pass with the change. It also reports test-backend-ops -b CUDA0 passing. The PR is still open as of September 24, so these should be read as upstream PR validation, not as proof that the fix is already in a released llama.cpp build.
The broader lesson: graph identity includes lifetime
Inference runtimes aggressively reuse everything they can:
CUDA graphs
KV blocks
workspaces
allocation plans
compiled kernels
prefix states
That reuse is where much of the performance comes from.
But every cache needs a correct key.
For a graph allocator, shape alone is not a complete identity when input/output flags change how long buffers must live.
The neat part of #29370 is that the fix makes that hidden assumption explicit.
A tensor is not only its dimensions and bytes.
It also has a lifetime.
And if the lifetime changes, the old memory plan may no longer be the same plan at all.