llama.cpp Unknown Model Architecture 'mllama': What to Do
Fix the llama.cpp error unknown model architecture 'mllama' by recognizing that current llama.cpp does not implement the Mllama vision architecture and choosing a supported path.
Approximately 4 min read
If llama.cpp stops with:
error loading model architecture: unknown model architecture: 'mllama'
the important fact is that current llama.cpp does not implement an mllama model architecture path.
That means this is not a VRAM problem and not something you can fix with a larger context, different GPU layers, another quant level, or an mmproj file.
The model metadata says:
general.architecture = mllama
but the runtime has no matching architecture implementation to dispatch.
The short answer
For an Mllama / Llama 3.2 Vision GGUF that reports general.architecture = mllama:
- confirm the runtime you are actually launching;
- update once to current llama.cpp to rule out an old binary;
- if current master still reports
unknown model architecture: 'mllama', stop changing memory settings; - use a runtime that supports that model architecture, or use a llama.cpp-supported model instead.
As of September 30, 2026, current llama.cpp source does not list an LLM_ARCH_MLLAMA architecture and its Hugging Face converter does not contain an Mllama conversion handler.
Why mmproj does not fix this error
A multimodal launch can contain two different artifacts:
language model GGUF
vision/projector artifact
The projector cannot compensate for an unsupported language-model architecture.
Upstream issue #14951 showed a Llama 3.2 11B Vision GGUF plus an mmproj file. llama.cpp successfully opened the GGUF metadata and printed:
general.architecture = mllama
then stopped at:
unknown model architecture: 'mllama'
The failure therefore happened at architecture dispatch, before a projector could turn unsupported model code into a supported path.
First, make sure you are not running an old binary
Record the exact executable:
command -v llama-cli
readlink -f "$(command -v llama-cli)"
llama-cli --version
If you launch through a GUI, container, Python package, service unit, or wrapper, verify which bundled llama.cpp revision it actually uses.
Updating is still a useful first check because architecture support changes quickly. But for this specific architecture, current master still has no Mllama model entry, so repeatedly rebuilding the same current code is not a fix.
Confirm the GGUF architecture
Use llama.cpp metadata output or another GGUF inspector and look for:
general.architecture = mllama
If the file instead says:
general.architecture = llama
then you have a different problem and should diagnose the first loader error from that file.
Do not rename the metadata from mllama to llama. The two architectures are not interchangeable merely because both belong to the Llama family.
Why changing quantization does not help
The loader rejects the architecture before normal tensor execution.
So trying:
Q4_K_M -> Q5_K_M
Q5_K_M -> Q8_0
Q8_0 -> F16
will not add missing runtime architecture code.
A different quant can matter after an architecture is supported and tensor loading begins. It cannot solve an unknown architecture identifier.
Why context and GPU flags do not help
These options operate later in initialization:
--ctx-size
--gpu-layers
--split-mode
--tensor-split
--cache-type-k
--cache-type-v
If architecture dispatch fails first, the runtime never reaches the stage where those settings can solve anything.
The same applies to freeing VRAM. More free memory does not make an unknown architecture known.
What are the practical options?
If you need the vision model
Use a runtime that explicitly supports the Mllama architecture and the vision path for the exact checkpoint you are using.
Verify support by model class and version rather than by a generic claim such as “supports Llama 3.2.” Text-only Llama support does not automatically mean Llama 3.2 Vision / Mllama support.
If you only need text generation
Use a text model whose GGUF architecture is already supported by your llama.cpp build rather than forcing the Mllama vision checkpoint through the wrong loader.
If you need llama.cpp specifically
Track upstream architecture support and test again only when the source actually contains an Mllama implementation. Do not treat a newer build number by itself as proof of support.
A useful verification check
After switching model or runtime, verify the failure stage moved past architecture dispatch.
For llama.cpp, a supported model should proceed from metadata loading into concrete model/tensor loading and context initialization rather than ending at:
unknown model architecture
Then run a small prompt before tuning context or offload settings.
Similar-looking errors need different fixes
unknown model architecture: 'mllama'
means the runtime has no architecture implementation for the identifier in the GGUF.
tensor '...' not found
means architecture dispatch succeeded but the runtime expected a tensor that is absent.
LLAMA_SPLIT_MODE_TENSOR not implemented for architecture 'qwen4exp'
means the model architecture is recognized, but a specific multi-GPU split strategy is unsupported.
failed to allocate buffer for kv cache
means model loading got much farther and a runtime memory allocation failed.
For the broader decision tree, see llama.cpp Errors and Fixes.
What to include in a bug report
Record:
llama.cpp version / commit
exact executable path
OS and backend
model repository and revision
exact GGUF filename
general.architecture value
whether an mmproj is being used
exact command
first loader error
For this exact Search Console error, the key diagnostic is simple: general.architecture = mllama reaches a llama.cpp build that has no Mllama architecture implementation. Change the supported runtime/model path, not memory tuning knobs.