Troubleshooting

llama.cpp unknown model architecture 'glm5next': Cause and Fix

Diagnose llama.cpp's unknown model architecture 'glm5next' error, verify the runtime you are actually using, and check whether upstream support has landed.

Approximately 7 min read

If llama.cpp stops with:

llama_model_load: error loading model: unknown model architecture: 'glm5next'

the first thing to check is runtime compatibility, not the quantization level.

The GGUF says its architecture is glm5next, but the llama.cpp runtime that is actually opening the file does not recognize that architecture identifier. Changing the same model from Q4 to Q5 or Q8 normally does not fix an architecture-dispatch failure, because the loader has failed before tensor precision is the central question.

As of September 26, 2026, upstream issue ggml-org/llama.cpp #27922 is still open. The two relevant implementation PRs, #27773 and #27754, are also open rather than merged. Current upstream source inspected for this article still maps unrecognized architecture strings to LLM_ARCH_UNKNOWN and throws the exact unknown model architecture error.

So the correct fix is not “upgrade to latest and it will definitely work.” First establish what runtime you are using and whether support is present in that runtime.

What the error actually means

A GGUF contains metadata describing the model, including its architecture. llama.cpp reads that architecture name and resolves it to one of the architectures implemented by the runtime.

Current llama.cpp source follows this path conceptually:

GGUF general.architecture
        |
        v
architecture string lookup
        |
        +-- known -> create the matching model implementation
        |
        +-- unknown -> LLM_ARCH_UNKNOWN
                      -> throw "unknown model architecture: '...'"

For glm5next, the failure means the running loader has no registered implementation for that identifier.

That is different from successfully recognizing the architecture and then discovering that a tensor is missing, malformed, the wrong shape, or too large for available memory.

Why this usually is not a bad Q4 quant

Quantization changes how model tensors are represented. It does not normally change the model-family identifier in the GGUF metadata.

If these files all declare the same architecture:

GLM-5.3-Flash-Q4_K_M.gguf  -> glm5next
GLM-5.3-Flash-Q5_K_M.gguf  -> glm5next
GLM-5.3-Flash-Q8_0.gguf    -> glm5next

then a runtime that does not know glm5next will reject all three at the same dispatch step.

A different quant can matter later if the runtime supports the architecture but a particular tensor type, backend kernel, converter output, or quantization path is incompatible. That produces a different diagnosis.

Do not spend hours redownloading larger quants until you have proved the runtime recognizes the architecture.

Check the runtime you are actually executing

Start with the executable path and version:

command -v llama-cli
readlink -f "$(command -v llama-cli)"
llama-cli --version

If you built from a llama.cpp checkout, also record its commit:

git -C /path/to/llama.cpp rev-parse HEAD

If you are already in the checkout:

git rev-parse HEAD
./build/bin/llama-cli --version

Then compare the binary you invoke with the tree you just updated.

A common failure pattern is:

~/src/llama.cpp      -> freshly pulled source
/usr/local/bin/llama-cli -> older installed binary actually in PATH

Rebuilding one does not update the other.

Frontends, containers and Python wrappers can bundle another llama.cpp

This error often survives an apparently successful upgrade because the user upgraded the wrong copy.

A desktop frontend, Docker image, Python package, server bundle, or another local-AI application may ship its own llama.cpp revision. Your standalone ~/llama.cpp checkout can be current while the application that opens the GGUF is weeks behind.

For a container, record the image reference and inspect the binary inside that container rather than on the host.

For a Python binding, record the package version and how it was built. A system llama-cli version does not prove which llama.cpp revision the Python extension contains.

For a GUI, use its own backend/runtime information if exposed. Do not assume “app version 1.2.3” maps to current llama.cpp master unless the vendor documents that mapping.

Confirm whether your source tree knows glm5next

If you have the source used to build the runtime, a direct check is useful:

git grep -n '"glm5next"' -- src

No match in the architecture registration/model implementation path is strong evidence that this source tree cannot load a GGUF whose architecture is glm5next.

Also record:

git status --short
git branch --show-current
git rev-parse HEAD

This matters if you are testing an unmerged feature branch. “My llama.cpp is new” is not enough; an unmerged model-support PR can contain code that current master does not.

The converter/runtime timing gap

New model families can create a temporary split between two parts of the ecosystem:

  1. a converter or model publisher can emit a GGUF using a new architecture name;
  2. the stable or mainline runtime that a user has installed may not yet implement that architecture.

That produces a valid-looking GGUF which the runtime still cannot instantiate.

The reverse can also happen: a runtime branch supports an architecture, while a GGUF was converted with older or incompatible assumptions.

Record both sides:

runtime:
  llama.cpp commit/version
  frontend or wrapper version
  backend/build options

artifact:
  model repository
  exact GGUF filename
  model revision
  converter or quantizer provenance if known

This pair is much more useful than “latest llama.cpp + GLM Q4.”

Check upstream support before rebuilding again

For this exact architecture, the current upstream tracking points are:

At publication time, these support PRs are not merged. That means updating an ordinary master checkout cannot provide code that is not on master.

If one of those PRs later merges, the next questions are:

  1. Which commit contains the merge?
  2. Does the binary you are running include that commit?
  3. Has your package/frontend/container released a build containing it?

A GitHub PR being “approved,” “ready,” or producing working test binaries is not the same as being part of the runtime you installed.

Minimal reproduction

Strip the launch down to the model and a small context so unrelated tuning flags do not hide the classification:

./build/bin/llama-cli   -m /path/to/model.gguf   -c 512   -n 1   -p "test"

If the same unknown model architecture: 'glm5next' appears immediately, context size, KV cache size and sampling settings are not your first problem.

Preserve the complete log from the first loader error onward.

Fixes, in order

1. Identify the actual binary

Resolve PATH, GUI backend, container image, or Python binding first. Do not rebuild a source checkout that is not being executed.

2. Compare its commit with upstream support status

If your binary predates merged support, update the actual runtime.

If support is still only in open PRs, ordinary master does not yet contain it.

3. Rebuild cleanly after support has landed

Once support is confirmed in the branch you intend to run, use a clean or known build directory so stale objects do not blur the result. Follow the current upstream build instructions for your backend rather than copying an old CMake command from a forum post.

Then verify again:

./build/bin/llama-cli --version
git rev-parse HEAD

4. Re-test the same GGUF before changing quants

Keep the artifact constant for the first compatibility test. If the architecture error disappears and a new tensor/backend error appears, that is progress: the loader has moved to a later stage.

5. Only then investigate conversion or quantization compatibility

Once the runtime recognizes glm5next, a tensor-name, shape, missing-op, backend, or quantization failure becomes meaningful on its own terms.

How to verify the fix

A successful fix is not “the error string changed.”

Verify that:

  1. the runtime log identifies the expected llama.cpp version/commit;
  2. glm5next is accepted rather than rejected during architecture dispatch;
  3. the model proceeds through tensor loading and context initialization;
  4. a minimal prompt produces output without a later fatal error.

If the architecture error disappears but loading stops on a missing tensor or wrong shape, diagnose that new error separately.

Similar-looking errors that need different fixes

Error What it usually means First place to look
unknown model architecture: 'glm5next' runtime has no architecture mapping/implementation runtime version and upstream support
tensor '...' not found runtime expects a tensor that the GGUF does not contain converter/model/runtime compatibility
tensor '...' has wrong shape tensor dimensions disagree with runtime expectation model config, converter and runtime revisions
failed to allocate buffer for kv cache backend could not reserve the required KV buffer context, concurrency, KV type, device memory
backend CUDA/Vulkan/ROCm/Metal error device/backend failed a specific operation deepest backend error and device placement

The broad llama.cpp errors and fixes guide covers the classification tree. For KV allocation specifically, use the KV-cache allocation troubleshooting hub.

What to include in an upstream bug report

Include enough information for another person to reproduce the same loader path:

llama.cpp commit/version:
binary path:
frontend/wrapper version:
OS:
backend:
GPU(s):
model repository:
model revision:
exact GGUF filename(s):
converter/quantizer provenance:
exact command:
first meaningful error:
full relevant log:
tested current master?:
tested an unmerged support PR?:

For glm5next, the most important field today is still the exact runtime commit. Until support is merged, two people both saying “latest llama.cpp” may actually be testing different branches with fundamentally different architecture support.

Sources and further reading

Continue reading