LongCat Sparse Lands in llama.cpp — But the Attention Path Is Still Dense
A source-path experiment on the first LongCat-Flash-Lite-Sparse llama.cpp fork: LSA token selection is real, but selected tokens are converted into a mask over full K/V rather than compact sparse attention execution.
Approximately 6 min read
A community llama.cpp fork now loads LongCat-Flash-Lite-Sparse, the 69B-A3B LongCat variant that replaces dense MLA with LongCat Sparse Attention (LSA) and advertises native context lengths up to 1 million tokens. Upstream llama.cpp does not currently provide this architecture path, so the release is interesting for a reason that has little to do with the uncensored checkpoint attached to it: somebody had to map a new sparse-attention model into llama.cpp.
We did a quick source-path experiment before deciding whether a large GGUF benchmark was worth the download.
The result is more nuanced than either “just model plumbing” or “llama.cpp now has a native sparse-attention backend.”
The fork implements the LSA selection semantics, including a real top-K token-selection path. But in the attention path we traced, those selected indices are converted into a sparse mask over the full KV history. The code then passes full K/V tensors into the ordinary MHA builder. It does not gather a compact top-K K/V working set before attention.
That distinction matters enormously for the 1M-context claim.
The Quick Experiment
Instead of downloading tens of gigabytes of weights, we traced one falsifiable question through the fork:
After LSA selects its tokens, does llama.cpp actually shorten K/V before the attention operation?
The LongCat model code exposes the relevant metadata directly. The sparse path expects an indexer head size of 128, a top-K budget of 2048, 16 initial tokens, 1024 local tokens, and a CLI reuse factor of 2. It builds the LongCat indexer Q/K path and produces a top_k tensor once the KV history crosses the sparse threshold.
The model then calls the DSA-flavored build_attn() overload and passes that top_k tensor into the graph builder.
So far, this is real sparse-attention architecture work rather than a loader-only port.
The decisive behavior appears one layer lower in src/llama-graph.cpp.
When top_k is present, the graph creates a mask initialized to negative infinity, writes zero into rows identified by the selected indices, combines that with the ordinary causal mask, and then retrieves K from the KV cache. V is a view over that same full cached tensor. Finally, it calls the normal multi-head attention builder with those full K/V tensors and the newly constructed mask.
Conceptually, the checked path is:
LSA indexer
-> top-K token indices
-> full-size attention mask
selected rows = allowed
other rows = -INF
-> full cached K
-> full cached V
-> ordinary MHA
What we did not find in that path is:
LSA indices
-> gather selected K rows
-> gather selected V rows
-> compact ~K-token attention
That gives us a clean implementation result without requiring a 69B runtime test.
This Is More Than Model Plumbing
Calling the fork “just plumbing” would be unfair.
The LongCat path has architecture-specific LSA metadata checks, indexer Q/K construction, top-K selection, owner/reuse behavior across blocks, a sparse activation threshold, and integration with llama.cpp’s DSA-oriented KV infrastructure. The LongCat model graph explicitly switches to a DSA attention input path for the sparse architecture and threads the selected indices into attention.
In other words, the implementation knows which historical tokens LongCat wants to attend to.
That is the semantic layer of sparse attention.
The important caveat is that semantic sparsity and computational sparsity are not the same thing.
Three Different Meanings of “Sparse Attention”
This fork is a useful example of why sparse-attention support should be split into three layers:
| Layer | Status in the checked fork |
|---|---|
| Architecture/model plumbing | Implemented |
| LSA sparse selection semantics | Implemented |
| Compact sparse attention execution | Not observed in the traced path |
The third layer is where the biggest long-context performance claim would come from.
If the model selects roughly 2K useful historical positions out of a 128K or 1M context, an ideal sparse execution path would avoid doing attention work over the entire history. That normally means gathering, paging, or otherwise exposing only the selected K/V subset to the attention kernel.
Here, the checked graph still materializes the selection as a mask and supplies the full K/V history to ordinary MHA.
That can preserve the model’s intended attention semantics while failing to realize the full computational advantage of sparse execution.
The DSA Cache Is Not the New Part
One subtlety is worth correcting before this turns into a misleading headline.
The fork already contained llama.cpp’s DSA KV-cache infrastructure on the earlier dense LongCat branch. The sparse LongCat branch reuses that machinery rather than introducing an entirely new LongCat-specific sparse-cache subsystem from scratch.
The genuinely LongCat-specific work we found sits mainly in the LSA indexer, top-K selection/reuse rules, model metadata/tensors, and the conversion of those selections into the DSA attention path.
So the interesting implementation claim is not “a brand-new sparse KV allocator landed.”
It is:
LongCat’s sparse selection policy has been expressed inside llama.cpp, but the selected set is currently enforced through masking rather than compact K/V execution in the path we traced.
What 1M Context Means Here
LongCat-Flash-Lite-Sparse genuinely declares a 1M native context architecture. The official model describes LSA as extending DeepSeek Sparse Attention with streaming-aware indexing and a bounded selection budget.
But “the architecture supports 1M context” and “llama.cpp executes 1M context at roughly a 2K attention working-set cost” are two separate claims.
The first is a model/configuration property.
The second is a runtime property that still needs measurement.
The current source trace makes the second claim doubtful enough that downloading the full model solely to demonstrate a 1M context launch is not yet the best experiment.
The Benchmark That Would Settle It
The next useful test is deliberately smaller than 1M:
4K -> 16K -> 64K -> 128K context
At each point we would record:
- prompt-processing throughput;
- token-generation throughput;
- VRAM and system-RAM usage;
- KV-cache size;
- the LSA selected-token count; and
- if possible, the actual K dimension reaching the attention kernel.
The falsifiable prediction is simple.
If execution is truly compact after sparse selection, attention cost should track the selected working set much more closely than total context length after the sparse threshold is crossed.
If prompt cost and attention memory continue scaling strongly with full context length while selected K remains near 2048, then the fork is sparse in semantics but still substantially dense in execution.
There is no reason to begin with a 1M-token run. A 4K-to-128K scaling curve should reveal the direction quickly.
Why This Port Still Matters
This is exactly the kind of community work that makes llama.cpp interesting.
New architectures frequently arrive before inference engines have first-class support. Someone has to decode the reference implementation, map the checkpoint metadata, reproduce indexing behavior, preserve cache semantics, and make the model run inside an engine whose abstractions were designed before that model existed.
The LongCat fork appears to have done a meaningful portion of that work.
It also exposes the next engineering problem very clearly: turning sparse token selection into sparse execution.
That is a better research target than simply reporting that a 69B GGUF launches.
For RAMGPT, the next step is therefore not a generic quality benchmark. It is a scaling experiment designed to answer one implementation question:
Does LongCat’s 2048-token sparse selection actually bound attention work in llama.cpp, or does the current masked-full-KV path still scale with total context?
That result will tell us whether the fork is already a practical 1M-context sparse runtime — or the correct semantic foundation for one.