llama.cpp Tensor 'Wrong Shape' or 'Not Found': How to Fix It
Fix llama.cpp check_tensor_dims failures by separating missing tensors from wrong shapes, then checking model metadata, converter provenance, runtime support, and MTP assumptions.
Approximately 5 min read
If llama.cpp stops with check_tensor_dims: tensor '...' not found or tensor '...' has wrong shape, do not start by changing quant level, context length, or GPU memory settings. These are structural model-loading errors: the runtime expects a named tensor with a specific shape and the GGUF does not satisfy that contract.
Use this order:
- Confirm the exact GGUF and all required shards.
- Record the llama.cpp build actually opening the file.
- Separate not found from wrong shape.
- Compare model config, converter provenance, and runtime revision.
- Test a fresh F16/BF16 conversion before blaming quantization.
- Re-quantize only after the unquantized conversion loads.
The broad llama.cpp errors and fixes guide covers the whole loader tree. This page focuses only on tensor-contract failures.
What check_tensor_dims means
Current llama.cpp loader code first looks up the requested tensor. If a required tensor is absent it throws a tensor '...' not found error. If the tensor exists but its dimensions disagree with the dimensions expected by the model implementation, it throws tensor '...' has wrong shape; expected ..., got ....
That gives you the first classification:
not found
-> runtime expected a tensor name the GGUF does not contain
wrong shape
-> tensor name exists, but its dimensions disagree with runtime expectations
Neither error automatically means “bad Q4.”
Missing tensor: inspect the model contract
Upstream issue #26916 is a useful recent example. A Qwen3.5-Hybrid model converted successfully to F16 and then quantized successfully, but llama.cpp failed while loading it with:
check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found
The report showed regular transformer blocks through blk.31 while GGUF metadata also advertised a NextN/MTP-related prediction layer. The reporter suspected a mismatch between converter output and what the runtime expected for that auxiliary layer.
The broader lesson is more important than that one checkpoint: successful conversion and quantization do not prove that the generated GGUF matches the loader’s current architecture contract.
A missing tensor can come from a converter/runtime mismatch, metadata that causes the loader to expect extra layers, a model derivative that differs from the supported base architecture, or an incomplete artifact. Do not rename an unrelated tensor to make the error disappear.
Wrong shape: the name matched, the dimensions did not
A Search Console query reaching RAMGPT contains this exact family:
tensor 'output.weight' has wrong shape;
expected 5120, 248320, got 5120, 32768, 1, 1
Those numbers prove that the runtime and file disagree about output.weight; they do not prove a single universal cause. For an output matrix, this can be consistent with a vocabulary/output-size or model-definition mismatch, but check the original model config, tokenizer/config files, converter revision, and architecture implementation before assigning the cause.
Historical upstream reports show the same error family arising from different model-definition and runtime-support gaps. Treat the expected and actual shapes as evidence, not as a diagnosis by themselves.
Keep converter and runtime provenance together
Record both sides:
runtime:
llama.cpp version/build
commit if built from source
wrapper/frontend version
backend
artifact:
original model repo and revision
exact GGUF filename(s)
converter revision
quantizer revision
config/tokenizer revision
A common failure is converting with one recent llama.cpp checkout and loading with an older runtime bundled inside a GUI, container, Python binding, or another application. The reverse can also happen.
“Latest” is not a reproducible version. Record a release or commit.
Test the unquantized conversion first
If you control the source model, isolate conversion from quantization:
original HF checkpoint
|
v
fresh converter
|
v
F16/BF16 GGUF
|
+-- same tensor error -> config/converter/runtime contract
|
+-- loads -> quantize and test again
If the F16/BF16 GGUF already fails on the same tensor, comparing Q4_K_M, Q5_K_M and Q8 files is wasted effort. Stay focused on architecture support, config metadata, and converter/runtime compatibility.
MTP and NextN metadata are part of compatibility
Newer architectures can carry auxiliary prediction layers for speculative decoding. If the failing tensor is just beyond the normal block range, inspect MTP/NextN metadata and verify that the runtime and converted artifact agree about whether those tensors exist.
Do not invent missing MTP weights. Either the source model, converter metadata, or runtime expectation must change so the structural contract becomes consistent.
Similar errors need different fixes
unknown model architecture fails earlier: the runtime cannot dispatch the model family. A check_tensor_dims error means architecture handling got far enough to request concrete tensors.
Memory failures are different again. Context size, KV type, GPU layers, and batch size can solve allocation errors; they do not normally make a required tensor appear or change a stored tensor shape.
For failed to allocate buffer for kv cache, use the KV-cache allocation guide. For failed to allocate compute pp buffers, use the dedicated compute-buffer guide.
Minimal reproduction
Use the smallest normal load you can and preserve the complete first loader failure:
llama-cli -m /path/to/model.gguf -c 512 -n 1 -p "test"
Record the llama.cpp version/commit, binary path, OS/backend, model repository and revision, exact GGUF files, converter/quantizer revisions, first failing tensor, and expected versus actual shape.
How to verify the fix
A real fix should pass the stage that previously failed. Confirm that the intended runtime is executing, tensor loading completes, context initialization completes, and a minimal prompt generates output.
If F16/BF16 loads but the derived quant does not, the quantization/artifact path becomes relevant. If both fail on the same tensor, remain focused on model definition, metadata, converter, and runtime compatibility.
The useful mental model is simple: a tensor-contract error means the file and loader disagree about model structure. Diagnose that contract before touching performance settings.