Qwen3.8 MTP 'output_hc_norm.weight not found' in llama.cpp
If a Qwen3.8-Flash-Next MTP draft GGUF fails on output_hc_norm.weight, separate target-model loading from detached draft-head compatibility, inspect the sidecar tensor layout, and verify the exact llama.cpp MTP implementation.
Approximately 4 min read
A Qwen3.8-Flash-Next target model can load normally in llama.cpp and then fail only when an MTP draft GGUF is added:
llama_model_load: error loading model:
check_tensor_dims: tensor 'output_hc_norm.weight' not found
That is a different problem from the target model being unsupported.
The high-value distinction is:
target model loads
+
draft sidecar fails
=
investigate MTP head layout and loader contract
Do not start by shrinking context or changing GPU layers.
Reproduce the target without MTP
First remove the draft model and speculative flags.
Prove that the main model can initialize and generate.
If it does, record:
target only: PASS
target + MTP sidecar: FAIL
That immediately narrows the investigation.
The missing tensor is being requested while the draft model is loading, not while the target’s normal graph is initializing.
Why this error has appeared
Upstream issue #29174 reports Qwen3.8-Flash-Next MTP draft files failing on current builds with either:
output_hc_norm.weight not found
or, for another layout:
token_embd.weight not found
The target model itself runs. The incompatibility is in the draft-head path.
Qwen3.8 MTP support has gone through several implementations and proposed loader changes. Some sidecars contain only the detached draft head. Others include or expect additional shared tensors.
A loader that assumes a self-contained model can therefore ask a head-only file for tensors that are intentionally absent.
The reverse can also happen: a sidecar can use a naming or block-indexing convention that a particular loader revision does not expect.
A CLI flag is not proof of a complete architecture path
Seeing draft-mtp in the command-line help proves the CLI knows the option.
It does not prove that every architecture and every community MTP sidecar layout is supported by that build.
The compatibility unit is closer to:
this llama.cpp revision
+ this model architecture
+ this draft-head format
+ this converter/export convention
Inspect the sidecar before changing runtime settings
Answer these questions from the GGUF metadata:
Does output_hc_norm.weight exist?
Does token_embd.weight exist?
Are hyper-connection tensors indexed at blk.0 or a later block?
Is the file documented as head-only or self-contained?
A missing tensor name and a present-but-differently-indexed tensor are different failures.
Do not fix either one by blindly renaming tensors. Tensor names encode the runtime’s model contract.
Watch absolute versus relative block indexing
Discussion around upstream MTP work documented sidecars whose hyper-connection tensors were indexed at their absolute position in the full model, while a loader path expected relative indexing for a detached draft.
For example, a sidecar can contain names like:
blk.48.hc_...
while a loader asks for:
blk.0.hc_...
That mismatch is structural. Extra VRAM does not fix it.
Match the sidecar to its documented revision
For experimental speculative decoding, a model card or publishing discussion should ideally name:
- llama.cpp commit or PR;
- exporter/converter revision;
- whether the head is detached or grafted;
- expected main-model family;
- required flags.
If that provenance is absent, treat the sidecar as an experiment rather than a plug-and-play accessory.
Why changing quantization usually does not help
The failure is a required-name lookup:
tensor 'output_hc_norm.weight' not found
A different Q4/Q5 level does not create a tensor the file never contained.
Quantization can matter when a conversion pipeline drops or rewrites tensors, but the correct fix is then to repair the conversion/layout, not to keep trying random bit rates.
A clean diagnostic matrix
Use the smallest matrix that isolates the contract:
A. target only
B. target + published MTP sidecar
C. target + alternate documented sidecar layout
D. same pair on the exact llama.cpp revision named by the sidecar author
Record success/failure and the first meaningful error.
If A passes and B fails on a missing head tensor, the target GGUF should no longer be the primary suspect.
Do not mix performance debugging with loader debugging
MTP raises a second question after it loads:
Does speculative decoding actually improve tokens/s on this hardware?
That is a benchmark question.
output_hc_norm.weight not found is earlier than that.
Keep the stages separate:
1. file/loader contract
2. graph initialization
3. functional generation
4. acceptance behavior
5. throughput
There is no value benchmarking acceptance rate until stage 1 passes.
Repro template
llama.cpp build/commit:
target model repo + exact GGUF:
draft model repo + exact GGUF:
target-only command/result:
MTP command/result:
draft GGUF tensor-name evidence:
converter/export provenance:
OS/backend/GPU:
first meaningful loader error:
This is much more actionable than “MTP doesn’t work.”
Bottom line
For Qwen3.8-Flash-Next, output_hc_norm.weight not found should be treated as a draft-head compatibility signal.
Prove the target works alone, inspect the sidecar’s actual tensor layout, and match it to the exact llama.cpp MTP implementation it was exported for.
The missing-tensor error is telling you that two components disagree about the draft-model contract. Solve that contract first; tune performance later.