Troubleshooting

Does llama.cpp Support NVFP4? Check the Exact Path Before You Convert

If you searched whether llama.cpp supports NVFP4, the answer depends on the checkpoint, converter, model architecture, backend, and quant-scale path. Use this checklist to identify the real blocker.

Approximately 6 min read

The short answer to “does llama.cpp support NVFP4?” is:

parts of the stack do, but that does not mean every NVFP4 checkpoint can be converted and run on every backend.

Current llama.cpp master defines an NVFP4 GGML tensor type, and the current conversion code contains explicit handling for NVFP4-style quantization metadata. That is real support.

But a working end-to-end path also needs the checkpoint format, model architecture, conversion logic, scale metadata, backend kernels, and hardware path to agree.

So when an NVFP4 model fails, first identify which layer rejected it instead of treating “NVFP4 support” as one binary feature.

Start by separating five questions

Use this checklist:

1. Does the converter recognize the source checkpoint's NVFP4 format?
2. Can it represent the needed tensors and scales in GGUF/GGML?
3. Does the model architecture map those tensors correctly?
4. Does the selected backend implement the required operations?
5. Does your hardware use the intended accelerated path?

A yes at one layer does not prove the next layer works.

Current master does recognize NVFP4

In current ggml/include/ggml.h, llama.cpp defines:

GGML_TYPE_NVFP4

The current conversion layer also contains explicit NVFP4 detection for compressed-tensors metadata and per-layer quantization information.

That means an old answer such as “llama.cpp has no NVFP4 type at all” is no longer accurate.

If your error comes from a recent build, look deeper.

Conversion support can still be the blocker

A Hugging Face checkpoint may use NVFP4 but encode its quantization metadata differently from another NVFP4 checkpoint.

Current conversion code checks quantization metadata such as:

quant_method
quant_format
quant_algo
config groups
per-layer quantization entries

So two checkpoints that both advertise NVFP4 can still follow different conversion paths.

If conversion fails, record the exact repository and inspect its config.json, quantization config, and tensor names before changing runtime flags.

Do not assume a CUDA flag can fix a converter that never produced a valid GGUF.

Quantization scales matter

NVFP4 is not just “4-bit weights.”

The quantized values depend on scale information. Upstream discussion #22042 explains why NVFP4 needs clear handling of both low-precision blocks and the scaling information required to interpret them correctly.

That matters when troubleshooting because a runtime can recognize the tensor type but still lack the exact scale semantics needed for a particular checkpoint path.

This is also why upstream PR #28898 is important. It is still draft work aimed at making FP8/NVFP4 quantization-scale handling more first-class in GGML and improving direct conversion of scaled checkpoints.

So the practical state is:

NVFP4 type exists
!=
every NVFP4 safetensors checkpoint has a stable direct-conversion path

Architecture support is separate from datatype support

A model loader does more than read tensor dtypes.

It maps named tensors into architecture-specific roles:

attention projections
MLP projections
experts
normalization
output head
recurrent state
other architecture-specific tensors

If that mapping does not understand the checkpoint’s NVFP4-related tensors or scale tensors, loading can fail even though GGML_TYPE_NVFP4 exists.

This is why an error about missing tensors, unexpected tensor counts, or an unsupported architecture should be diagnosed as a loader/mapping problem first.

Do not immediately re-quantize the model.

Backend support is another boundary

After conversion and loading, the selected backend still has to execute the operations required by the model.

A build may therefore behave differently across:

CUDA
Vulkan
Metal
CPU
other backend combinations

If an NVFP4 GGUF loads but fails when computation begins, capture:

llama-cli --version
llama-cli --list-devices

Then record the backend, GPU model, driver/runtime version, and the first failing operation.

That is more useful than reporting only “NVFP4 does not work.”

Blackwell hardware does not make every path automatic

NVFP4 is especially associated with NVIDIA Blackwell hardware, but hardware capability alone does not guarantee that your exact llama.cpp execution path uses a native optimized kernel.

There are at least three separate questions:

Can the model representation be loaded?
Can the backend execute it correctly?
Is the accelerated hardware path actually selected?

A fallback or partially supported path can still run differently from the path you expected.

Measure the actual runtime rather than inferring support from the GPU name.

A practical troubleshooting sequence

If you are trying to run an NVFP4 model, keep the test small.

1. Update and record the exact revision

llama-cli --version

For a source checkout:

git rev-parse HEAD

NVFP4 work is moving quickly enough that a months-old answer may no longer match master.

2. Test conversion separately

Run the converter before debugging GPU flags.

If conversion fails, save the complete first traceback and the checkpoint’s quantization metadata.

Classify the failure:

unknown quantization format
unsupported tensor mapping
missing scale tensor
unexpected tensor name
unsupported architecture

3. Validate the GGUF before performance tuning

If conversion succeeds, inspect the produced file and confirm that the runtime recognizes the model and tensor types.

Do not begin with large context, many GPU layers, or concurrency.

4. Run the smallest inference

Use one prompt, one sequence, and a modest context.

If the model loads and generates correctly, then move to performance tuning.

If it fails, the first backend or tensor error is the evidence you need.

5. Change one variable at a time

Do not simultaneously change:

llama.cpp revision
GGUF
backend
context
GPU offload
batch size
KV type

Otherwise you will not know which change fixed the path.

How to read common failure classes

Converter rejects the checkpoint

Likely layer:

checkpoint metadata
converter
scale representation
architecture mapping

Start there.

GGUF exists but loader reports missing or wrong tensors

Likely layer:

architecture mapping
conversion output
tensor naming/layout

This is not primarily a CUDA tuning problem.

Model loads but backend reports an unsupported operation

Likely layer:

backend implementation
kernel path
device capability

Try to reproduce with the same GGUF on another supported backend if possible.

Model runs but performance is unexpectedly poor

Likely layer:

fallback path
kernel selection
GPU offload
prompt-processing path
batching

At that point it is a performance investigation, not an NVFP4-format detection problem.

Why older search results can be misleading

NVFP4 support in llama.cpp has changed over time.

Older discussions correctly described a period when the type or conversion path was missing. Current master now contains explicit NVFP4 definitions and conversion logic.

At the same time, open work around quantization scales shows that “support exists” should not be simplified into “all NVFP4 checkpoints are universally supported.”

Both statements can be true:

llama.cpp has real NVFP4 support
and
some NVFP4 checkpoint/backend combinations can still fail

The exact revision and failure stage decide which one matters to you.

Bottom line

If your question is “does llama.cpp support NVFP4?”, do not stop at yes or no.

Current master recognizes NVFP4 in GGML and in conversion logic, but end-to-end compatibility still depends on the checkpoint’s scale representation, model architecture, converter path, backend, and hardware.

For a failing model, capture the exact revision and the first error, then classify the failure as conversion, loader mapping, backend execution, or performance.

That will usually tell you what actually needs to change.

Sources and further reading

Continue reading