Local AI

llama.cpp Cuts Qwen4exp Indexer Score Memory: What Actually Changed

An October 3 llama.cpp change rewrites Qwen4exp indexer scoring around the lightning indexer, reducing live score-buffer memory and extending backend handling across CUDA, Metal, and Vulkan.

Approximately 5 min read

On October 3, llama.cpp merged a Qwen4exp change with a deceptively simple headline: halve the indexer score memory.

The important part is how it gets there. The previous path could keep two large floating-point score tensors alive at once. The new implementation reformulates the work around llama.cpp’s lightning-indexer operation so it does not need to materialize the same per-head score structure in the old way.

This matters most at long context, where an intermediate tensor that scales with token count can become a real part of the memory budget.

The old live-buffer shape

The upstream commit describes the old scoring path as holding two tensors shaped approximately like:

[n_pool, n_idx_h, n_tokens]

in F32. One product held raw scores and another held rectified scores.

That means the live memory for this stage grows with:

number of pooled keys
x indexer heads
x active tokens
x 4 bytes

At short context the cost can be easy to ignore. At long context it becomes exactly the kind of graph-local allocation that can turn an otherwise fitting configuration into a marginal one.

The commit explicitly calls these among the largest graph buffers at long context.

This is a compute-buffer optimization

Nothing here says model weights became smaller. The optimization attacks temporary compute state.

A useful runtime budget is:

weights
+ KV cache / recurrent state
+ compute buffers
+ graph temporaries
+ allocator headroom
= practical memory footprint

A model can keep the same GGUF, the same quantization and the same KV-cache type while still needing less peak memory because one graph stage has been rewritten.

This is why file size alone is a poor proxy for whether a long-context configuration fits.

Why the lightning indexer is the cleaner formulation

The commit history shows an intermediate implementation that processed heads separately, then a review-driven simplification: the required reduction can be expressed as the lightning indexer with uniform head weights.

Conceptually, the score is still built from rectified per-head similarity terms, but the specialized operator can consume the pooled keys and head weights without exposing the large per-head score tensor as a persistent graph value.

The systems lesson is broader than Qwen4exp:

If a reduction is mathematically expressible as one specialized operator, materializing every intermediate can be unnecessary memory traffic.

CUDA: four-head support was added

The patch extends CUDA’s lightning indexer to a four-head case. It chooses a vector kernel because four heads are too few for the WMMA tile path used for larger supported head counts.

That is not cosmetic. Once the high-level graph depends on a specialized operator, each backend has to accept the exact shapes that graph can produce.

A graph rewrite that saves memory on paper is useless if one backend rejects the new operator shape.

Metal: head count becomes a function constant

The Metal implementation previously carried a fixed head-count assumption in this path.

The patch moves the head count into a Metal function constant and zero-fills the partial final head tile. The graph can therefore be more general while the compiled GPU pipeline remains specialized to the actual head count.

That pattern appears throughout high-performance inference:

general model graph
-> runtime shape
-> backend specialization
-> cached specialized kernel

Vulkan: the dispatch changed more substantially

The Vulkan portion changes the work decomposition so a workgroup scores a tile of keys against multiple tokens. Keys are staged into shared memory and reused rather than driving the earlier flat/subgroup dispatch.

The new dispatch is expressed in terms of key and token tiles.

This is relevant to memory traffic as well as temporary allocation, but the upstream commit states that compute-buffer size and speed are unchanged for the reviewed implementation. RAMGPT therefore does not claim an end-to-end speedup from the source change alone.

What “halve the indexer score memory” does not mean

It does not mean:

The claim is narrower: the indexer scoring stage no longer needs the previous pair of large score buffers live together.

The practical benefit depends on how large that stage was relative to the rest of the runtime.

Why long-context users should care

At long context, several memory terms grow at once:

fixed model weights
+ context-dependent KV/state
+ context-dependent graph buffers
+ allocator headroom

If a graph temporary scales with token count, reducing it can move the failure boundary even when the KV cache is unchanged.

This also explains why “free VRAM before launch” is not enough information. Peak graph allocation can happen later during context creation or evaluation.

How to verify it locally

A useful A/B test needs two llama.cpp revisions: one immediately before the patch and one including commit 889edf43.

Keep fixed:

same GGUF
same context
same batch / ubatch
same KV types
same split mode
same GPU
same prompt
same output length

Record:

peak device memory
context initialization success/failure
prefill throughput
decode throughput

For a long-context Qwen4exp workload, peak memory is the primary metric. Throughput is secondary because the upstream claim is specifically about score-buffer memory.

The deeper runtime lesson

Local-AI memory tuning is no longer just about quantizing weights.

Modern runtimes increasingly win by:

That work can change the practical hardware envelope without changing a single model parameter.

For Qwen4exp users, commit 889edf43 is worth tracking for a precise reason: it removes a long-context-sensitive intermediate from the critical memory path.

The correct follow-up is measurement, not hype.

Sources and further reading

Continue reading