llama.cpp MTP Support: Use draft-mtp and Diagnose Failures
Current llama.cpp supports MTP speculative decoding through draft-mtp. Learn how to verify model support, enable it, isolate failed MTP contexts, and measure whether it actually improves your workload.
Approximately 6 min read
Yes: current upstream llama.cpp supports Multi-Token Prediction (MTP) as a speculative-decoding mode.
The important flag is:
--spec-type draft-mtp
Current upstream documentation describes draft-mtp as using MTP heads from the main model.
That does not mean every model can use it.
The real requirement is:
runtime supports draft-mtp
+
model artifact contains compatible MTP data
+
the selected frontend actually exposes the option
If any of those three is missing, turning on MTP can fail or do nothing useful.
First confirm your binary knows draft-mtp
Check the binary you are actually running:
llama-server --help | grep -i -E 'spec|mtp'
or:
llama-cli --help | grep -i -E 'spec|mtp'
You want to see draft-mtp among the supported speculative modes.
If the option is absent, do not conclude that llama.cpp as a project lacks MTP support.
You may simply be using an older binary, a downstream fork, a packaged build behind upstream, or a frontend that does not expose the current option.
Version the executable before debugging the model.
Establish a baseline without MTP
Start with a normal load.
For example:
llama-server -m /models/model.gguf
Confirm the model loads, prompt completion works, the expected chat template works, context allocation succeeds, and memory use is understood.
Only then add MTP.
That prevents an ordinary model-loading problem from being misclassified as an MTP problem.
Enable MTP explicitly
For a supported model and current binary, the relevant speculative mode is:
llama-server \
-m /models/model.gguf \
--spec-type draft-mtp
Use the exact option names shown by your binary’s --help because downstream wrappers can expose different subsets of upstream flags.
A config field is not proof that the tensors exist
Some Hugging Face model configurations declare MTP or NextN layers.
That is useful metadata, but it is not sufficient evidence that a converted artifact contains every tensor the runtime needs.
The correct verification chain is:
config declares MTP
-> source tensor index contains MTP/NextN tensors
-> converter preserves them
-> GGUF metadata identifies them correctly
-> runtime creates the MTP context
If the chain breaks, the model can still be perfectly usable without MTP.
Do not invent missing MTP tensors or copy unrelated weights into their place.
If you see “failed to create MTP context”
Treat that as a separate layer from ordinary generation.
A useful isolation sequence is:
1. load model without speculative decoding
2. confirm normal inference works
3. enable only draft-mtp
4. remove any external draft model
5. retry with a short context
6. capture exact runtime version and startup log
This matters because upstream issue reports have shown failure modes around MTP context creation and combinations of speculative modes.
The smallest reproducible configuration tells you whether the failure belongs to the checkpoint, converted GGUF, runtime version, speculative configuration, or an interaction between modes.
Do not enable MTP globally for every model
A router or shared preset may apply the same speculative options to many models.
That is convenient until one of those models does not contain compatible MTP data.
Prefer per-model capability:
model A -> draft-mtp enabled
model B -> no MTP
model C -> external draft model
instead of one global flag applied indiscriminately.
A model that lacks MTP should still be allowed to run normally.
MTP support and MTP speedup are different questions
A model can support MTP correctly without producing a useful speedup on every prompt.
Speculative decoding depends on acceptance.
Conceptually:
MTP proposes extra tokens
-> main computation validates them
-> accepted tokens save sequential decode steps
-> rejected tokens reduce the benefit
So performance depends on model, prompt, sampling, context length, GPU, runtime, MTP implementation, and acceptance behavior.
There is no honest universal statement such as “MTP = 2x faster” without measurements.
Measure baseline and MTP under the same conditions
A useful A/B test keeps these fixed:
same model file
same runtime commit
same prompt
same context
same sampling
same output limit
same GPU placement
same KV-cache type
Then compare:
baseline decode tok/s
MTP decode tok/s
time to first token
accepted draft tokens
rejected draft tokens
peak VRAM
output correctness
If the MTP run is faster but changes sampling or context settings, the comparison is not controlled.
Long context can change the result
MTP may add state or change the runtime’s memory behavior.
At short context, that overhead may be trivial.
At long context, KV cache and other buffers may dominate available VRAM.
So test more than one context point.
For a consumer GPU, a useful progression is 1K-4K, 8K, 16K, and 32K, then higher if memory remains stable.
The configured maximum context is not the same as the practical context for one GPU and one cache format.
Frontends can lag upstream
A common source of confusion is that llama.cpp supports a feature before a downstream application exposes it.
If the frontend has no MTP control, check its bundled llama.cpp revision, command-line pass-through support, preset schema, and server launch arguments.
Do not assume the absence of a checkbox means the backend project lacks the feature.
Likewise, an API wrapper may expose generation while hiding speculative-decoding controls.
A useful failure map
--spec-type unknown
-> binary/frontend is behind or does not expose current upstream option
normal load fails
-> fix the base model load first
normal load works, draft-mtp fails to initialize
-> inspect model MTP capability, GGUF conversion, and runtime version
draft-mtp runs but speed is unchanged
-> inspect acceptance, workload, context, and GPU bottleneck
draft-mtp is slower
-> compare accepted tokens and overhead; disable it for that workload
one model works, another fails under the same preset
-> make MTP configuration model-specific
Do not confuse MTP with an external draft model
Both belong to speculative decoding, but they are not the same mechanism.
An external draft model uses a separate smaller model to propose tokens.
MTP uses prediction heads associated with the main model.
That distinction matters when reading examples or issue reports.
A failure involving an external draft checkpoint does not automatically prove the main model’s MTP path is broken.
What I would record in a bug report
If MTP fails, include:
llama.cpp commit or build version
exact model and GGUF
quantization
GPU and backend
full launch command
context size
KV-cache type
whether normal inference works
whether external draft is also enabled
the first MTP-specific error
That is much more actionable than “MTP doesn’t work.”
The practical answer
For current upstream llama.cpp, the answer to “does llama.cpp support MTP?” is yes.
Use the current draft-mtp speculative mode, but treat checkpoint capability as a separate requirement.
The safe workflow is:
update runtime
-> confirm normal model load
-> verify MTP data exists
-> enable draft-mtp only
-> compare controlled baseline
-> keep it only if correctness and performance are both acceptable
Support is the first gate.
A measured benefit on your actual model and hardware is the second.