llama.cpp SPLIT_AXIS_UNKNOWN: Why --split-mode tensor Can Abort
Diagnose GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN) by separating tensor-split graph incompatibility from ordinary VRAM problems, then testing layer split, KV types, model paths and revisions.
Approximately 5 min read
If llama.cpp exits with:
GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN) failed
you are not looking at an ordinary out-of-VRAM message.
The assertion comes from the graph’s split-state analysis. The runtime is trying to determine how a tensor or operation is distributed across devices. For a node whose input split states cannot be reconciled, the result becomes SPLIT_AXIS_UNKNOWN, and the current meta-backend path asserts rather than continuing.
This has repeatedly appeared with:
--split-mode tensor
on multi-GPU setups.
First test: switch only the split mode
Keep the same model and approximate device ratio but replace tensor split with layer split.
If this fails:
-sm tensor
-ts 1,1
and this starts:
-sm layer
-ts 1,1
you have strong evidence that the problem belongs to tensor-split graph placement rather than a corrupt GGUF or simple aggregate-memory shortage.
Issue #27964 gives this kind of control: the reported Qwen3.8-Flash-Next configuration aborted under tensor split, while layer split with the same 1,1 ratio served requests.
That does not prove every SPLIT_AXIS_UNKNOWN report has the same root cause, but it is a high-value first test.
What the assertion means in current source
The meta backend tracks a split state for tensors.
Very roughly, a value can be:
split on a tensor dimension
mirrored
partial
unknown
For generic operations, llama.cpp examines the source split states. If they disagree in a way the current handler cannot reconcile, the code assigns GGML_BACKEND_SPLIT_AXIS_UNKNOWN.
The assertion means:
This graph operation reached a combination of split states the current code path does not know how to represent safely.
That is more specific than “multi-GPU failed.”
Why model architecture matters
Different architectures create different graphs.
A dense transformer, gated-attention model, recurrent hybrid and speculative-draft graph do not present the splitter with identical operations.
Issue #26902 is instructive. The reporter traced a failure to a multiply involving an attention gate whose split granularity differed from the attention output path.
The proposed local change made the states agree for that model, but the report also explains why forcing that rule globally could be unsafe.
That is why users should not paste a one-line source hack from an issue into every build.
A split-state workaround can be shape- and architecture-specific.
KV-cache types can expose another path
Issue #27116 reproduced the assertion with:
--split-mode tensor
--cache-type-k iq4_nl
--cache-type-v iq4_nl
on two RTX 5060 Ti GPUs.
That does not mean IQ4_NL is universally broken.
It means the diagnostic matrix should include cache type as an independent variable:
tensor split + current KV types
tensor split + f16/q8 KV types
layer split + current KV types
layer split + f16/q8 KV types
Change one dimension at a time.
Remove unrelated complexity
Before reporting the issue, remove features that are not required to trigger it:
- speculative decoding;
- draft models;
- unusual offload overrides;
- very large context;
- custom tensor overrides;
- server middleware.
Issue #27964 is useful because the reporter showed a failure without DFlash2, MTP or speculative decoding. That narrowed the problem to tensor split itself for that model path.
Record the exact build
Do not write:
latest llama.cpp
Write:
llama-server --version
git rev-parse HEAD
Issue #25829 is explicitly framed as a regression after a particular build boundary. Whether that exact regression applies to your architecture is a separate question, but the methodological lesson is durable: preserve a known-good SHA when tensor split breaks.
Split mode and tensor-split ratio are different
A ratio such as 1,1 does not prove that the tensor strategy itself is valid for the graph.
Conceptually:
split mode
= what distribution strategy is used
tensor split ratio
= how much is assigned to each device
You can have a sensible device ratio and still hit an unsupported tensor-level split state.
Practical decision tree
SPLIT_AXIS_UNKNOWN
|
+-- multi-GPU tensor split enabled?
| +-- no -> inspect graph/backend-specific issue
| +-- yes
| +-- layer split works?
| | +-- yes -> tensor-split path is primary suspect
| +-- simpler KV types work?
| | +-- yes -> cache/split interaction is implicated
| +-- older known-good commit works?
| +-- yes -> regression window is bounded
What not to do
Do not immediately:
- redownload a large GGUF with no corruption evidence;
- reduce context repeatedly and call that a fix;
- assume total free VRAM explains an internal split-state assertion;
- apply a model-specific split-granularity patch to unrelated architectures;
- report only the final assert line without split mode, GPU list and model.
The complete command matters.
A strong upstream report
Include:
llama.cpp commit/build:
OS:
backend:
GPU models:
model repo + exact GGUF:
exact command:
split mode:
tensor split ratio:
KV types:
layer-split result:
single-GPU result if practical:
known-good commit:
full relevant log:
For this class of problem, the most valuable comparison is often:
same model
same GPUs
same build
tensor split = fail
layer split = pass
That converts a generic crash into evidence about the graph-placement layer.
Bottom line
SPLIT_AXIS_UNKNOWN is a graph-distribution failure signal.
Treat it as a split-state compatibility problem first, not as a generic memory error.
Layer split is the cleanest control, build SHA is essential provenance, and architecture-specific fixes should stay architecture-specific until upstream code establishes a general rule.