Local AI

llama.cpp Can Pick the Wrong GGUF From a Model Directory: What PR #28809 Changes

A source-level look at llama.cpp PR #28809, which makes server router discovery prefer a GGUF whose filename matches its model directory instead of relying on directory iteration order.

Approximately 4 min read

A directory full of valid GGUF files looks harmless until software has to decide which one is the model.

That is the small but operationally important problem behind llama.cpp PR #28809. The proposed fix changes server router discovery so that, when a model directory contains several GGUF files, it prefers a file named exactly after the directory: <directory>.gguf.

This is not a new architecture or a CUDA optimization. It is a model-selection bug, and that makes it especially useful to understand: the files can all be readable GGUFs while the server still chooses the wrong one.

The failure mode

According to the upstream PR, router discovery can currently select the wrong main model when one directory contains multiple GGUF files. The selected file can depend on directory iteration order.

Imagine this layout:

models/
  my-model/
    my-model.gguf
    draft.gguf
    mmproj.gguf
    alternate.gguf

A human sees my-model.gguf as the obvious primary model. A generic directory scan does not automatically know that.

The dangerous part is that this is not necessarily a clean “file not found” failure. The router may discover a GGUF file successfully; it is simply not the GGUF the operator intended as the main model.

That distinction matters for automation. A deployment can have the correct directory, correct permissions, valid GGUF files, and still resolve the wrong artifact.

What PR #28809 proposes

The proposed rule is deliberately narrow:

model directory: foo/
preferred main file: foo/foo.gguf

The filename comparison is case-sensitive.

The PR says existing behavior remains in place for draft-model and mmproj detection, first-shard precedence, and fallback selection. In other words, the directory-name match is a preference for identifying the main model, not a wholesale rewrite of GGUF discovery.

This is a good compatibility property. A directory that does not contain an exact <directory>.gguf match can still use the previous fallback behavior.

Why directory iteration order is a bad model selector

Filesystem enumeration is an implementation detail, not model metadata.

If selection effectively becomes “first plausible GGUF encountered,” the result can vary with directory contents and ordering. Adding an auxiliary GGUF can therefore change behavior without changing the launch command.

That creates an unpleasant class of incident:

same command
+ one additional GGUF artifact
= different model selected

For local experiments this is confusing. For a server that discovers models automatically, it is a reproducibility problem.

A naming convention gives the operator a deterministic way to express intent without requiring a new manifest format.

Sharded models are the edge case to protect

A simplistic fix could accidentally break sharded GGUF discovery. The upstream author explicitly calls out preservation of first-shard precedence and says regression coverage includes sharded models.

That matters because a sharded model is not represented by one ordinary <directory>.gguf file. Selection logic must continue recognizing the shard set correctly rather than treating the new naming preference as universal.

The proposed hierarchy is therefore more subtle than “always load the matching filename”: preserve specialized detection rules, prefer the exact directory-name match where applicable, then retain fallback behavior.

What the tests tell us

The PR reports regression tests for three relevant layouts:

The author states that the regression test fails before the fix and passes afterward, and that all nine stages of local CPU CI completed successfully.

Those are upstream author-reported validation results, not RAMGPT measurements. PR #28809 remains an upstream change under review as of September 14, 2026, so operators should not assume every installed llama.cpp build already contains this behavior.

Practical diagnosis

If llama-server or router-based discovery appears to load an unexpected model from a directory containing several GGUF files, inspect the directory before blaming quantization or model metadata.

Ask four questions:

  1. Are several .gguf files present in the same model directory?
  2. Which one is intended to be the main model?
  3. Does its basename exactly match the directory name?
  4. Is the llama.cpp build old enough that PR #28809 is not present?

For a directory named Qwen-Lab, the proposed deterministic filename is:

Qwen-Lab/Qwen-Lab.gguf

not qwen-lab.gguf; the proposed match is case-sensitive.

Until the change is merged and reaches the build you use, explicit model paths remain the least ambiguous option when directory discovery is not required.

The broader engineering lesson

Model-loading reliability is not only about parsing tensors and recognizing architecture identifiers. Discovery logic is part of the runtime contract too.

As local inference directories accumulate main models, draft models, multimodal projectors, adapters, and alternate quantizations, “find a GGUF” stops being a sufficient rule. The runtime needs deterministic semantics for deciding what role each artifact plays.

PR #28809 addresses one narrow ambiguity with a convention humans already tend to expect: the file named after the model directory is probably the main model.

That is a small change, but it removes filesystem ordering from a decision that should be reproducible.

Sources and further reading

Continue reading