Troubleshooting

llama.cpp SPLIT_AXIS_UNKNOWN: Why --split-mode tensor Can Abort

Diagnose GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN) by separating tensor-split graph incompatibility from ordinary VRAM problems, then testing layer split, KV types, model paths and revisions.

Approximately 5 min read

If llama.cpp exits with:

GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN) failed

you are not looking at an ordinary out-of-VRAM message.

The assertion comes from the graph’s split-state analysis. The runtime is trying to determine how a tensor or operation is distributed across devices. For a node whose input split states cannot be reconciled, the result becomes SPLIT_AXIS_UNKNOWN, and the current meta-backend path asserts rather than continuing.

This has repeatedly appeared with:

--split-mode tensor

on multi-GPU setups.

First test: switch only the split mode

Keep the same model and approximate device ratio but replace tensor split with layer split.

If this fails:

-sm tensor
-ts 1,1

and this starts:

-sm layer
-ts 1,1

you have strong evidence that the problem belongs to tensor-split graph placement rather than a corrupt GGUF or simple aggregate-memory shortage.

Issue #27964 gives this kind of control: the reported Qwen3.8-Flash-Next configuration aborted under tensor split, while layer split with the same 1,1 ratio served requests.

That does not prove every SPLIT_AXIS_UNKNOWN report has the same root cause, but it is a high-value first test.

What the assertion means in current source

The meta backend tracks a split state for tensors.

Very roughly, a value can be:

split on a tensor dimension
mirrored
partial
unknown

For generic operations, llama.cpp examines the source split states. If they disagree in a way the current handler cannot reconcile, the code assigns GGML_BACKEND_SPLIT_AXIS_UNKNOWN.

The assertion means:

This graph operation reached a combination of split states the current code path does not know how to represent safely.

That is more specific than “multi-GPU failed.”

Why model architecture matters

Different architectures create different graphs.

A dense transformer, gated-attention model, recurrent hybrid and speculative-draft graph do not present the splitter with identical operations.

Issue #26902 is instructive. The reporter traced a failure to a multiply involving an attention gate whose split granularity differed from the attention output path.

The proposed local change made the states agree for that model, but the report also explains why forcing that rule globally could be unsafe.

That is why users should not paste a one-line source hack from an issue into every build.

A split-state workaround can be shape- and architecture-specific.

KV-cache types can expose another path

Issue #27116 reproduced the assertion with:

--split-mode tensor
--cache-type-k iq4_nl
--cache-type-v iq4_nl

on two RTX 5060 Ti GPUs.

That does not mean IQ4_NL is universally broken.

It means the diagnostic matrix should include cache type as an independent variable:

tensor split + current KV types
tensor split + f16/q8 KV types
layer split + current KV types
layer split + f16/q8 KV types

Change one dimension at a time.

Remove unrelated complexity

Before reporting the issue, remove features that are not required to trigger it:

Issue #27964 is useful because the reporter showed a failure without DFlash2, MTP or speculative decoding. That narrowed the problem to tensor split itself for that model path.

Record the exact build

Do not write:

latest llama.cpp

Write:

llama-server --version
git rev-parse HEAD

Issue #25829 is explicitly framed as a regression after a particular build boundary. Whether that exact regression applies to your architecture is a separate question, but the methodological lesson is durable: preserve a known-good SHA when tensor split breaks.

Split mode and tensor-split ratio are different

A ratio such as 1,1 does not prove that the tensor strategy itself is valid for the graph.

Conceptually:

split mode
= what distribution strategy is used

tensor split ratio
= how much is assigned to each device

You can have a sensible device ratio and still hit an unsupported tensor-level split state.

Practical decision tree

SPLIT_AXIS_UNKNOWN
|
+-- multi-GPU tensor split enabled?
|   +-- no -> inspect graph/backend-specific issue
|   +-- yes
|       +-- layer split works?
|       |   +-- yes -> tensor-split path is primary suspect
|       +-- simpler KV types work?
|       |   +-- yes -> cache/split interaction is implicated
|       +-- older known-good commit works?
|           +-- yes -> regression window is bounded

What not to do

Do not immediately:

The complete command matters.

A strong upstream report

Include:

llama.cpp commit/build:
OS:
backend:
GPU models:
model repo + exact GGUF:
exact command:
split mode:
tensor split ratio:
KV types:
layer-split result:
single-GPU result if practical:
known-good commit:
full relevant log:

For this class of problem, the most valuable comparison is often:

same model
same GPUs
same build
tensor split = fail
layer split = pass

That converts a generic crash into evidence about the graph-placement layer.

Bottom line

SPLIT_AXIS_UNKNOWN is a graph-distribution failure signal.

Treat it as a split-state compatibility problem first, not as a generic memory error.

Layer split is the cleanest control, build SHA is essential provenance, and architecture-specific fixes should stay architecture-specific until upstream code establishes a general rule.

Sources and further reading

Continue reading