llama.cpp Qwen4Exp: Split-Mode Tensor Is Not Implemented
Fix the llama.cpp error 'LLAMA_SPLIT_MODE_TENSOR not implemented for architecture qwen4exp' by using a supported split mode and understanding the architecture gate.
Approximately 5 min read
If llama.cpp exits with:
LLAMA_SPLIT_MODE_TENSOR not implemented for architecture 'qwen4exp'
the useful answer is short:
Stop using
--split-mode tensorfor this architecture. Use the normal layer split instead.
This is not an out-of-memory error, not evidence that your GGUF is corrupt, and not something that -ts 1,1 can repair.
Current llama.cpp checks architecture support before model execution. If the requested split mode is tensor and the architecture does not advertise tensor-split support, model creation throws this exact error.
For Qwen3.8-Flash-Next files identified as qwen4exp, the practical multi-GPU path is currently layer splitting.
The fastest fix
If your command contains:
-sm tensor
or:
--split-mode tensor
change it to:
-sm layer
For example:
llama-server -m /models/qwen3.8-flash-next.gguf -ngl all -sm layer -ts 1,1 -c 32768
The exact tensor-split ratio is hardware-dependent, so -ts 1,1 is only an example for two similar GPUs.
The important change is:
tensor split -> layer split
Issue #27964 documents the same model family working with layer split while tensor split failed on a two-GPU CUDA system.
Why -ts does not make tensor mode supported
Two different options are easy to confuse:
-sm / --split-mode
-ts / --tensor-split
--split-mode selects the strategy.
Current CLI documentation describes the modes as:
none
layer
row
tensor
Layer is the normal multi-GPU mode. Tensor mode is experimental and is not implemented for every model architecture.
--tensor-split supplies proportions used to divide work or model data across devices. It does not add an implementation for a model architecture.
So this combination:
-sm tensor -ts 1,1
still requests the unsupported tensor strategy.
Changing only the ratio cannot bypass the architecture check.
The error happens deliberately in model creation
Current src/llama-model.cpp maps LLM_ARCH_QWEN4EXP to the Qwen4Exp model implementation.
Immediately after creating the model object, llama.cpp checks whether tensor split is supported.
The logic is conceptually:
if requested split mode is TENSOR
and architecture does not support TENSOR
throw "LLAMA_SPLIT_MODE_TENSOR not implemented for architecture ..."
That is much more informative than a random CUDA failure.
It means the runtime is refusing a configuration it does not know how to execute correctly.
This is why reducing context, changing KV-cache precision, or freeing VRAM is not the first fix.
The failure occurs before those memory-tuning techniques can solve the unsupported split strategy.
Why Qwen3.8-Flash-Next is not a simple dense Transformer
Qwen3.8-Flash-Next uses the qwen4exp architecture identifier in llama.cpp.
Its execution structure includes model-specific components beyond a plain stack of identical dense Transformer blocks.
Tensor splitting is not just “cut every matrix in half.”
A backend needs model-aware metadata describing which tensor axis can be partitioned and how the corresponding computation is reconstructed across devices.
If that metadata is missing or incomplete, forcing a tensor split can produce incorrect execution or assertions.
That is why current llama.cpp uses an explicit support gate rather than assuming every architecture can use the experimental mode.
An older symptom was SPLIT_AXIS_UNKNOWN
Issue #27964 was originally filed with a lower-level failure:
GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN) failed
The reporter used Qwen3.8-Flash-Next on two RTX 5090 GPUs with:
--split-mode tensor
--tensor-split 1,1
The same model and hardware worked with layer split.
Later reports in the same issue show the newer, clearer failure:
LLAMA_SPLIT_MODE_TENSOR not implemented for architecture 'qwen4exp'
From an operator’s perspective, the newer message is better.
Instead of allowing unsupported metadata to reach a later assertion, the runtime can reject the configuration at model creation.
Do not debug this as a KV-cache problem
A common reaction to any multi-GPU startup failure is to change:
context length
KV-cache type
batch size
micro-batch size
GPU layers
flash attention
Those are useful controls for allocation failures.
They are not the primary controls for an architecture capability error.
A simple diagnostic hierarchy is:
"not implemented for architecture"
-> configuration/support problem
"failed to allocate ..."
-> memory/capacity problem
CUDA illegal access / assertion during execution
-> backend/runtime correctness problem
Start with the most precise error, not the most familiar tuning knob.
A working layer split can still use both GPUs
Moving from tensor split to layer split does not mean falling back to one GPU.
With multiple CUDA devices and enough offloaded layers, llama.cpp can assign model layers across devices.
A typical two-GPU configuration may look like:
llama-server -m /models/model.gguf -ngl all -sm layer -ts 1,1
For unequal GPUs, adjust the proportions based on usable memory rather than assuming equal shares.
You may also need to leave headroom for KV cache and compute buffers.
Those are separate capacity questions after the split mode itself is valid.
Should you try row mode instead?
Do not treat another split mode as an automatic substitute merely because the parser accepts its name.
The upstream report establishes layer split as the known working path for this model family.
If your objective is simply to make Qwen3.8-Flash-Next load across multiple GPUs, use the documented working mode first.
Experiment with other modes only when the current llama.cpp revision explicitly supports the architecture and you have a reason to test them.
When tensor mode may become usable
The error describes current implementation support, not a permanent property of the model.
A future llama.cpp change could add the missing tensor-split metadata and execution support for qwen4exp.
So before assuming the limitation still exists months later, check:
current llama.cpp revision
current architecture support code
related upstream issue / PR status
But on a build that emits the exact architecture error, there is nothing useful to “force.”
The runtime has already told you that the requested strategy is not implemented.
Search-query answer
For the exact error:
llama_split_mode_tensor not implemented for architecture 'qwen4exp'
use:
-sm layer
instead of:
-sm tensor
Then tune -ts, context length, KV cache, batch sizes, and GPU allocation only after the model successfully enters a supported multi-GPU execution path.
That ordering avoids spending an hour on VRAM settings for a failure that is actually a feature gate.