Troubleshooting

llama.cpp Unknown Model Architecture 'mllama': What to Do

Fix the llama.cpp error unknown model architecture 'mllama' by recognizing that current llama.cpp does not implement the Mllama vision architecture and choosing a supported path.

Approximately 4 min read

If llama.cpp stops with:

error loading model architecture: unknown model architecture: 'mllama'

the important fact is that current llama.cpp does not implement an mllama model architecture path.

That means this is not a VRAM problem and not something you can fix with a larger context, different GPU layers, another quant level, or an mmproj file.

The model metadata says:

general.architecture = mllama

but the runtime has no matching architecture implementation to dispatch.

The short answer

For an Mllama / Llama 3.2 Vision GGUF that reports general.architecture = mllama:

  1. confirm the runtime you are actually launching;
  2. update once to current llama.cpp to rule out an old binary;
  3. if current master still reports unknown model architecture: 'mllama', stop changing memory settings;
  4. use a runtime that supports that model architecture, or use a llama.cpp-supported model instead.

As of September 30, 2026, current llama.cpp source does not list an LLM_ARCH_MLLAMA architecture and its Hugging Face converter does not contain an Mllama conversion handler.

Why mmproj does not fix this error

A multimodal launch can contain two different artifacts:

language model GGUF
vision/projector artifact

The projector cannot compensate for an unsupported language-model architecture.

Upstream issue #14951 showed a Llama 3.2 11B Vision GGUF plus an mmproj file. llama.cpp successfully opened the GGUF metadata and printed:

general.architecture = mllama

then stopped at:

unknown model architecture: 'mllama'

The failure therefore happened at architecture dispatch, before a projector could turn unsupported model code into a supported path.

First, make sure you are not running an old binary

Record the exact executable:

command -v llama-cli
readlink -f "$(command -v llama-cli)"
llama-cli --version

If you launch through a GUI, container, Python package, service unit, or wrapper, verify which bundled llama.cpp revision it actually uses.

Updating is still a useful first check because architecture support changes quickly. But for this specific architecture, current master still has no Mllama model entry, so repeatedly rebuilding the same current code is not a fix.

Confirm the GGUF architecture

Use llama.cpp metadata output or another GGUF inspector and look for:

general.architecture = mllama

If the file instead says:

general.architecture = llama

then you have a different problem and should diagnose the first loader error from that file.

Do not rename the metadata from mllama to llama. The two architectures are not interchangeable merely because both belong to the Llama family.

Why changing quantization does not help

The loader rejects the architecture before normal tensor execution.

So trying:

Q4_K_M -> Q5_K_M
Q5_K_M -> Q8_0
Q8_0 -> F16

will not add missing runtime architecture code.

A different quant can matter after an architecture is supported and tensor loading begins. It cannot solve an unknown architecture identifier.

Why context and GPU flags do not help

These options operate later in initialization:

--ctx-size
--gpu-layers
--split-mode
--tensor-split
--cache-type-k
--cache-type-v

If architecture dispatch fails first, the runtime never reaches the stage where those settings can solve anything.

The same applies to freeing VRAM. More free memory does not make an unknown architecture known.

What are the practical options?

If you need the vision model

Use a runtime that explicitly supports the Mllama architecture and the vision path for the exact checkpoint you are using.

Verify support by model class and version rather than by a generic claim such as “supports Llama 3.2.” Text-only Llama support does not automatically mean Llama 3.2 Vision / Mllama support.

If you only need text generation

Use a text model whose GGUF architecture is already supported by your llama.cpp build rather than forcing the Mllama vision checkpoint through the wrong loader.

If you need llama.cpp specifically

Track upstream architecture support and test again only when the source actually contains an Mllama implementation. Do not treat a newer build number by itself as proof of support.

A useful verification check

After switching model or runtime, verify the failure stage moved past architecture dispatch.

For llama.cpp, a supported model should proceed from metadata loading into concrete model/tensor loading and context initialization rather than ending at:

unknown model architecture

Then run a small prompt before tuning context or offload settings.

Similar-looking errors need different fixes

unknown model architecture: 'mllama'

means the runtime has no architecture implementation for the identifier in the GGUF.

tensor '...' not found

means architecture dispatch succeeded but the runtime expected a tensor that is absent.

LLAMA_SPLIT_MODE_TENSOR not implemented for architecture 'qwen4exp'

means the model architecture is recognized, but a specific multi-GPU split strategy is unsupported.

failed to allocate buffer for kv cache

means model loading got much farther and a runtime memory allocation failed.

For the broader decision tree, see llama.cpp Errors and Fixes.

What to include in a bug report

Record:

llama.cpp version / commit
exact executable path
OS and backend
model repository and revision
exact GGUF filename
general.architecture value
whether an mmproj is being used
exact command
first loader error

For this exact Search Console error, the key diagnostic is simple: general.architecture = mllama reaches a llama.cpp build that has no Mllama architecture implementation. Change the supported runtime/model path, not memory tuning knobs.

Sources and further reading

Continue reading