llama.cpp 'blk.32.attn_norm.weight' Not Found: Fix MTP Metadata
Fix the Qwen3.5-family blk.32.attn_norm.weight missing-tensor error when config metadata declares an MTP/NextN layer that the checkpoint does not contain.
Approximately 5 min read
If a Qwen3.5-family GGUF fails with:
check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found
check whether the source config declares an MTP/NextN layer that is not actually present in the checkpoint.
For that specific mismatch, the clean conversion workaround is:
python3 convert_hf_to_gguf.py /path/to/model --outtype bf16 --no-mtp --outfile model-bf16.gguf
Current llama.cpp accepts both --no-mtp and the alias --no-nextn for this purpose.
Do not start by lowering context, changing GPU layers, or trying another quant type. A missing blk.32 tensor is a model-structure mismatch, not a runtime-memory shortage.
The root cause seen upstream
llama.cpp issue #26916 documented this exact failure on a Qwen3.5 hybrid derivative.
The model had 32 normal transformer blocks, indexed:
blk.0
...
blk.31
but its metadata indicated one additional NextN/MTP prediction layer.
The generated GGUF therefore advertised a structure that caused the loader to expect another block. Loading reached:
blk.32.attn_norm.weight
but that tensor was not in the file.
An upstream maintainer summarized the issue directly: if a model does not include MTP layers, its config should say so; --no-mtp is the converter workaround.
Why conversion can succeed anyway
A conversion script can finish writing a GGUF even when metadata and the actual tensor inventory disagree in a way that only becomes fatal at load time.
The sequence can look like this:
HF checkpoint
-> conversion succeeds
-> quantization succeeds
-> llama.cpp load fails on blk.32
That does not prove the quantizer broke the model.
If the F16/BF16 GGUF already contains metadata that says an extra prediction layer exists while the corresponding tensors are absent, every derived quant can inherit the same structural problem.
Confirm it before reconverting
Inspect three things:
- the model config;
- the tensor names in the source checkpoint;
- the GGUF metadata and highest real block index.
The suspicious pattern is:
normal transformer layers: 32
highest real block: blk.31
MTP/NextN declared: 1
loader expects: blk.32.*
actual blk.32 tensors: none
On Qwen3.5-family models, the config key commonly involved is:
"mtp_num_hidden_layers": 1
The exact key and converter behavior can evolve, so use the current converter and inspect the actual model rather than copying metadata values blindly.
Use --no-mtp when the checkpoint has no MTP tensors
Current convert_hf_to_gguf.py documents:
--no-nextn, --no-mtp
as the option for excluding NextN speculative draft tensors from the converted GGUF.
For a derivative that inherited MTP metadata but did not ship MTP weights, convert the target model without the absent draft layer:
python3 convert_hf_to_gguf.py ./model --outtype bf16 --no-nextn --outfile ./model-bf16.gguf
Then test that unquantized GGUF before producing Q4, Q5, Q6, or Q8 variants.
Do not use --no-mtp blindly
Some models really do contain MTP tensors.
If the checkpoint contains a valid MTP head and you want speculative decoding, dropping it changes what you are publishing. Current llama.cpp can also export MTP separately with --mtp for supported model classes.
So the decision is:
config declares MTP
+ checkpoint contains matching MTP tensors
-> preserve/export them correctly
config declares MTP
+ checkpoint contains no MTP tensors
-> correct the config or convert target with --no-mtp
The flag is a workaround for an inconsistent artifact, not a universal Qwen3.5 optimization.
Why changing quantization usually does not help
Suppose BF16 conversion produces a GGUF whose metadata implies 33 blocks while the tensor inventory ends at block 31.
Quantizing that file to Q4_K_M changes representation and size. It does not manufacture a missing transformer or MTP block.
Use this test order:
fresh converter
-> BF16/F16 GGUF
-> load test
-> only then quantize
-> load test again
If BF16 fails on the same blk.32 tensor, stop testing different quant levels and fix the model contract first.
A minimal verification run
After reconversion:
llama-cli -m ./model-bf16.gguf -c 512 -n 16 -p "Hello"
You are looking for four checkpoints:
- tensor loading passes
blk.32without a missing-tensor exception; - model loading completes;
- context initialization completes;
- generation produces tokens.
Only after that should you quantize:
llama-quantize ./model-bf16.gguf ./model-q4_k_m.gguf Q4_K_M
Then repeat the same minimal load test on the quantized file.
If you downloaded someone else’s broken GGUF
The best fix is a corrected re-export from the publisher.
A user can sometimes patch GGUF metadata, but that is riskier because the correct values depend on the actual checkpoint. Do not blindly change block_count from 33 to 32 unless you have verified that the file really contains only 32 main blocks and no MTP tensors.
For a public model repo, report:
exact GGUF filename
llama.cpp version
qwen35 block_count / NextN metadata
highest blk.N index present
whether any MTP tensors exist
exact missing tensor
That gives the publisher enough evidence to regenerate the artifacts correctly.
Similar errors that are not this bug
A different missing tensor can have a different cause.
For example:
tensor 'output.weight' has wrong shape
means the tensor exists but its dimensions disagree with the loader.
unknown model architecture: 'qwen4exp'
means the runtime could not dispatch the architecture at all.
failed to allocate buffer for kv cache
is a memory or backend allocation path.
Use the broader wrong shape / not found guide when the missing tensor is not specifically the extra blk.32 pattern.
What to record for a reproducible bug report
Include:
source model repo and revision
config.json MTP/NextN fields
number of normal transformer layers
whether mtp.* or equivalent tensors exist
converter commit
converter command
GGUF metadata for block count / NextN
highest blk.N tensor present
quantizer commit and command
runtime commit
exact loader error
The diagnostic principle is straightforward: if metadata tells llama.cpp to load a 33rd block but the checkpoint only contains 32, fix the metadata/conversion contract instead of tuning runtime memory settings.