Does llama.cpp Support NVFP4? Check the Exact Path Before You Convert
If you searched whether llama.cpp supports NVFP4, the answer depends on the checkpoint, converter, model architecture, backend, and quant-scale path. Use this checklist to identify the real blocker.
Approximately 6 min read
The short answer to “does llama.cpp support NVFP4?” is:
parts of the stack do, but that does not mean every NVFP4 checkpoint can be converted and run on every backend.
Current llama.cpp master defines an NVFP4 GGML tensor type, and the current conversion code contains explicit handling for NVFP4-style quantization metadata. That is real support.
But a working end-to-end path also needs the checkpoint format, model architecture, conversion logic, scale metadata, backend kernels, and hardware path to agree.
So when an NVFP4 model fails, first identify which layer rejected it instead of treating “NVFP4 support” as one binary feature.
Start by separating five questions
Use this checklist:
1. Does the converter recognize the source checkpoint's NVFP4 format?
2. Can it represent the needed tensors and scales in GGUF/GGML?
3. Does the model architecture map those tensors correctly?
4. Does the selected backend implement the required operations?
5. Does your hardware use the intended accelerated path?
A yes at one layer does not prove the next layer works.
Current master does recognize NVFP4
In current ggml/include/ggml.h, llama.cpp defines:
GGML_TYPE_NVFP4
The current conversion layer also contains explicit NVFP4 detection for compressed-tensors metadata and per-layer quantization information.
That means an old answer such as “llama.cpp has no NVFP4 type at all” is no longer accurate.
If your error comes from a recent build, look deeper.
Conversion support can still be the blocker
A Hugging Face checkpoint may use NVFP4 but encode its quantization metadata differently from another NVFP4 checkpoint.
Current conversion code checks quantization metadata such as:
quant_method
quant_format
quant_algo
config groups
per-layer quantization entries
So two checkpoints that both advertise NVFP4 can still follow different conversion paths.
If conversion fails, record the exact repository and inspect its config.json, quantization config, and tensor names before changing runtime flags.
Do not assume a CUDA flag can fix a converter that never produced a valid GGUF.
Quantization scales matter
NVFP4 is not just “4-bit weights.”
The quantized values depend on scale information. Upstream discussion #22042 explains why NVFP4 needs clear handling of both low-precision blocks and the scaling information required to interpret them correctly.
That matters when troubleshooting because a runtime can recognize the tensor type but still lack the exact scale semantics needed for a particular checkpoint path.
This is also why upstream PR #28898 is important. It is still draft work aimed at making FP8/NVFP4 quantization-scale handling more first-class in GGML and improving direct conversion of scaled checkpoints.
So the practical state is:
NVFP4 type exists
!=
every NVFP4 safetensors checkpoint has a stable direct-conversion path
Architecture support is separate from datatype support
A model loader does more than read tensor dtypes.
It maps named tensors into architecture-specific roles:
attention projections
MLP projections
experts
normalization
output head
recurrent state
other architecture-specific tensors
If that mapping does not understand the checkpoint’s NVFP4-related tensors or scale tensors, loading can fail even though GGML_TYPE_NVFP4 exists.
This is why an error about missing tensors, unexpected tensor counts, or an unsupported architecture should be diagnosed as a loader/mapping problem first.
Do not immediately re-quantize the model.
Backend support is another boundary
After conversion and loading, the selected backend still has to execute the operations required by the model.
A build may therefore behave differently across:
CUDA
Vulkan
Metal
CPU
other backend combinations
If an NVFP4 GGUF loads but fails when computation begins, capture:
llama-cli --version
llama-cli --list-devices
Then record the backend, GPU model, driver/runtime version, and the first failing operation.
That is more useful than reporting only “NVFP4 does not work.”
Blackwell hardware does not make every path automatic
NVFP4 is especially associated with NVIDIA Blackwell hardware, but hardware capability alone does not guarantee that your exact llama.cpp execution path uses a native optimized kernel.
There are at least three separate questions:
Can the model representation be loaded?
Can the backend execute it correctly?
Is the accelerated hardware path actually selected?
A fallback or partially supported path can still run differently from the path you expected.
Measure the actual runtime rather than inferring support from the GPU name.
A practical troubleshooting sequence
If you are trying to run an NVFP4 model, keep the test small.
1. Update and record the exact revision
llama-cli --version
For a source checkout:
git rev-parse HEAD
NVFP4 work is moving quickly enough that a months-old answer may no longer match master.
2. Test conversion separately
Run the converter before debugging GPU flags.
If conversion fails, save the complete first traceback and the checkpoint’s quantization metadata.
Classify the failure:
unknown quantization format
unsupported tensor mapping
missing scale tensor
unexpected tensor name
unsupported architecture
3. Validate the GGUF before performance tuning
If conversion succeeds, inspect the produced file and confirm that the runtime recognizes the model and tensor types.
Do not begin with large context, many GPU layers, or concurrency.
4. Run the smallest inference
Use one prompt, one sequence, and a modest context.
If the model loads and generates correctly, then move to performance tuning.
If it fails, the first backend or tensor error is the evidence you need.
5. Change one variable at a time
Do not simultaneously change:
llama.cpp revision
GGUF
backend
context
GPU offload
batch size
KV type
Otherwise you will not know which change fixed the path.
How to read common failure classes
Converter rejects the checkpoint
Likely layer:
checkpoint metadata
converter
scale representation
architecture mapping
Start there.
GGUF exists but loader reports missing or wrong tensors
Likely layer:
architecture mapping
conversion output
tensor naming/layout
This is not primarily a CUDA tuning problem.
Model loads but backend reports an unsupported operation
Likely layer:
backend implementation
kernel path
device capability
Try to reproduce with the same GGUF on another supported backend if possible.
Model runs but performance is unexpectedly poor
Likely layer:
fallback path
kernel selection
GPU offload
prompt-processing path
batching
At that point it is a performance investigation, not an NVFP4-format detection problem.
Why older search results can be misleading
NVFP4 support in llama.cpp has changed over time.
Older discussions correctly described a period when the type or conversion path was missing. Current master now contains explicit NVFP4 definitions and conversion logic.
At the same time, open work around quantization scales shows that “support exists” should not be simplified into “all NVFP4 checkpoints are universally supported.”
Both statements can be true:
llama.cpp has real NVFP4 support
and
some NVFP4 checkpoint/backend combinations can still fail
The exact revision and failure stage decide which one matters to you.
Bottom line
If your question is “does llama.cpp support NVFP4?”, do not stop at yes or no.
Current master recognizes NVFP4 in GGML and in conversion logic, but end-to-end compatibility still depends on the checkpoint’s scale representation, model architecture, converter path, backend, and hardware.
For a failing model, capture the exact revision and the first error, then classify the failure as conversion, loader mapping, backend execution, or performance.
That will usually tell you what actually needs to change.