Troubleshooting

llama.cpp 'blk.32.attn_norm.weight' Not Found: Fix MTP Metadata

Fix the Qwen3.5-family blk.32.attn_norm.weight missing-tensor error when config metadata declares an MTP/NextN layer that the checkpoint does not contain.

Approximately 5 min read

If a Qwen3.5-family GGUF fails with:

check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found

check whether the source config declares an MTP/NextN layer that is not actually present in the checkpoint.

For that specific mismatch, the clean conversion workaround is:

python3 convert_hf_to_gguf.py   /path/to/model   --outtype bf16   --no-mtp   --outfile model-bf16.gguf

Current llama.cpp accepts both --no-mtp and the alias --no-nextn for this purpose.

Do not start by lowering context, changing GPU layers, or trying another quant type. A missing blk.32 tensor is a model-structure mismatch, not a runtime-memory shortage.

The root cause seen upstream

llama.cpp issue #26916 documented this exact failure on a Qwen3.5 hybrid derivative.

The model had 32 normal transformer blocks, indexed:

blk.0
...
blk.31

but its metadata indicated one additional NextN/MTP prediction layer.

The generated GGUF therefore advertised a structure that caused the loader to expect another block. Loading reached:

blk.32.attn_norm.weight

but that tensor was not in the file.

An upstream maintainer summarized the issue directly: if a model does not include MTP layers, its config should say so; --no-mtp is the converter workaround.

Why conversion can succeed anyway

A conversion script can finish writing a GGUF even when metadata and the actual tensor inventory disagree in a way that only becomes fatal at load time.

The sequence can look like this:

HF checkpoint
  -> conversion succeeds
  -> quantization succeeds
  -> llama.cpp load fails on blk.32

That does not prove the quantizer broke the model.

If the F16/BF16 GGUF already contains metadata that says an extra prediction layer exists while the corresponding tensors are absent, every derived quant can inherit the same structural problem.

Confirm it before reconverting

Inspect three things:

  1. the model config;
  2. the tensor names in the source checkpoint;
  3. the GGUF metadata and highest real block index.

The suspicious pattern is:

normal transformer layers: 32
highest real block:         blk.31
MTP/NextN declared:         1
loader expects:             blk.32.*
actual blk.32 tensors:      none

On Qwen3.5-family models, the config key commonly involved is:

"mtp_num_hidden_layers": 1

The exact key and converter behavior can evolve, so use the current converter and inspect the actual model rather than copying metadata values blindly.

Use --no-mtp when the checkpoint has no MTP tensors

Current convert_hf_to_gguf.py documents:

--no-nextn, --no-mtp

as the option for excluding NextN speculative draft tensors from the converted GGUF.

For a derivative that inherited MTP metadata but did not ship MTP weights, convert the target model without the absent draft layer:

python3 convert_hf_to_gguf.py   ./model   --outtype bf16   --no-nextn   --outfile ./model-bf16.gguf

Then test that unquantized GGUF before producing Q4, Q5, Q6, or Q8 variants.

Do not use --no-mtp blindly

Some models really do contain MTP tensors.

If the checkpoint contains a valid MTP head and you want speculative decoding, dropping it changes what you are publishing. Current llama.cpp can also export MTP separately with --mtp for supported model classes.

So the decision is:

config declares MTP
+ checkpoint contains matching MTP tensors
-> preserve/export them correctly

config declares MTP
+ checkpoint contains no MTP tensors
-> correct the config or convert target with --no-mtp

The flag is a workaround for an inconsistent artifact, not a universal Qwen3.5 optimization.

Why changing quantization usually does not help

Suppose BF16 conversion produces a GGUF whose metadata implies 33 blocks while the tensor inventory ends at block 31.

Quantizing that file to Q4_K_M changes representation and size. It does not manufacture a missing transformer or MTP block.

Use this test order:

fresh converter
  -> BF16/F16 GGUF
      -> load test
          -> only then quantize
              -> load test again

If BF16 fails on the same blk.32 tensor, stop testing different quant levels and fix the model contract first.

A minimal verification run

After reconversion:

llama-cli   -m ./model-bf16.gguf   -c 512   -n 16   -p "Hello"

You are looking for four checkpoints:

  1. tensor loading passes blk.32 without a missing-tensor exception;
  2. model loading completes;
  3. context initialization completes;
  4. generation produces tokens.

Only after that should you quantize:

llama-quantize   ./model-bf16.gguf   ./model-q4_k_m.gguf   Q4_K_M

Then repeat the same minimal load test on the quantized file.

If you downloaded someone else’s broken GGUF

The best fix is a corrected re-export from the publisher.

A user can sometimes patch GGUF metadata, but that is riskier because the correct values depend on the actual checkpoint. Do not blindly change block_count from 33 to 32 unless you have verified that the file really contains only 32 main blocks and no MTP tensors.

For a public model repo, report:

exact GGUF filename
llama.cpp version
qwen35 block_count / NextN metadata
highest blk.N index present
whether any MTP tensors exist
exact missing tensor

That gives the publisher enough evidence to regenerate the artifacts correctly.

Similar errors that are not this bug

A different missing tensor can have a different cause.

For example:

tensor 'output.weight' has wrong shape

means the tensor exists but its dimensions disagree with the loader.

unknown model architecture: 'qwen4exp'

means the runtime could not dispatch the architecture at all.

failed to allocate buffer for kv cache

is a memory or backend allocation path.

Use the broader wrong shape / not found guide when the missing tensor is not specifically the extra blk.32 pattern.

What to record for a reproducible bug report

Include:

source model repo and revision
config.json MTP/NextN fields
number of normal transformer layers
whether mtp.* or equivalent tensors exist
converter commit
converter command
GGUF metadata for block count / NextN
highest blk.N tensor present
quantizer commit and command
runtime commit
exact loader error

The diagnostic principle is straightforward: if metadata tells llama.cpp to load a 33rd block but the checkpoint only contains 32, fix the metadata/conversion contract instead of tuning runtime memory settings.

Sources and further reading

Continue reading