Troubleshooting

Qwen3.8 MTP 'output_hc_norm.weight not found' in llama.cpp

If a Qwen3.8-Flash-Next MTP draft GGUF fails on output_hc_norm.weight, separate target-model loading from detached draft-head compatibility, inspect the sidecar tensor layout, and verify the exact llama.cpp MTP implementation.

Approximately 4 min read

A Qwen3.8-Flash-Next target model can load normally in llama.cpp and then fail only when an MTP draft GGUF is added:

llama_model_load: error loading model:
check_tensor_dims: tensor 'output_hc_norm.weight' not found

That is a different problem from the target model being unsupported.

The high-value distinction is:

target model loads
+
draft sidecar fails
=
investigate MTP head layout and loader contract

Do not start by shrinking context or changing GPU layers.

Reproduce the target without MTP

First remove the draft model and speculative flags.

Prove that the main model can initialize and generate.

If it does, record:

target only: PASS
target + MTP sidecar: FAIL

That immediately narrows the investigation.

The missing tensor is being requested while the draft model is loading, not while the target’s normal graph is initializing.

Why this error has appeared

Upstream issue #29174 reports Qwen3.8-Flash-Next MTP draft files failing on current builds with either:

output_hc_norm.weight not found

or, for another layout:

token_embd.weight not found

The target model itself runs. The incompatibility is in the draft-head path.

Qwen3.8 MTP support has gone through several implementations and proposed loader changes. Some sidecars contain only the detached draft head. Others include or expect additional shared tensors.

A loader that assumes a self-contained model can therefore ask a head-only file for tensors that are intentionally absent.

The reverse can also happen: a sidecar can use a naming or block-indexing convention that a particular loader revision does not expect.

A CLI flag is not proof of a complete architecture path

Seeing draft-mtp in the command-line help proves the CLI knows the option.

It does not prove that every architecture and every community MTP sidecar layout is supported by that build.

The compatibility unit is closer to:

this llama.cpp revision
+ this model architecture
+ this draft-head format
+ this converter/export convention

Inspect the sidecar before changing runtime settings

Answer these questions from the GGUF metadata:

Does output_hc_norm.weight exist?
Does token_embd.weight exist?
Are hyper-connection tensors indexed at blk.0 or a later block?
Is the file documented as head-only or self-contained?

A missing tensor name and a present-but-differently-indexed tensor are different failures.

Do not fix either one by blindly renaming tensors. Tensor names encode the runtime’s model contract.

Watch absolute versus relative block indexing

Discussion around upstream MTP work documented sidecars whose hyper-connection tensors were indexed at their absolute position in the full model, while a loader path expected relative indexing for a detached draft.

For example, a sidecar can contain names like:

blk.48.hc_...

while a loader asks for:

blk.0.hc_...

That mismatch is structural. Extra VRAM does not fix it.

Match the sidecar to its documented revision

For experimental speculative decoding, a model card or publishing discussion should ideally name:

If that provenance is absent, treat the sidecar as an experiment rather than a plug-and-play accessory.

Why changing quantization usually does not help

The failure is a required-name lookup:

tensor 'output_hc_norm.weight' not found

A different Q4/Q5 level does not create a tensor the file never contained.

Quantization can matter when a conversion pipeline drops or rewrites tensors, but the correct fix is then to repair the conversion/layout, not to keep trying random bit rates.

A clean diagnostic matrix

Use the smallest matrix that isolates the contract:

A. target only
B. target + published MTP sidecar
C. target + alternate documented sidecar layout
D. same pair on the exact llama.cpp revision named by the sidecar author

Record success/failure and the first meaningful error.

If A passes and B fails on a missing head tensor, the target GGUF should no longer be the primary suspect.

Do not mix performance debugging with loader debugging

MTP raises a second question after it loads:

Does speculative decoding actually improve tokens/s on this hardware?

That is a benchmark question.

output_hc_norm.weight not found is earlier than that.

Keep the stages separate:

1. file/loader contract
2. graph initialization
3. functional generation
4. acceptance behavior
5. throughput

There is no value benchmarking acceptance rate until stage 1 passes.

Repro template

llama.cpp build/commit:
target model repo + exact GGUF:
draft model repo + exact GGUF:
target-only command/result:
MTP command/result:
draft GGUF tensor-name evidence:
converter/export provenance:
OS/backend/GPU:
first meaningful loader error:

This is much more actionable than “MTP doesn’t work.”

Bottom line

For Qwen3.8-Flash-Next, output_hc_norm.weight not found should be treated as a draft-head compatibility signal.

Prove the target works alone, inspect the sidecar’s actual tensor layout, and match it to the exact llama.cpp MTP implementation it was exported for.

The missing-tensor error is telling you that two components disagree about the draft-model contract. Solve that contract first; tune performance later.

Sources and further reading

Continue reading