Troubleshooting

llama.cpp File Not Found: Fix GGUF Paths, Shards, and Router 404s

Diagnose llama.cpp File Not Found errors by checking the actual GGUF path, process working directory, split-model shards, router configuration, and file readability before blaming model compatibility.

Approximately 5 min read

A llama.cpp File Not Found error is often simpler than it looks.

Before changing quantization, rebuilding CUDA, or blaming the model architecture, first prove that the process can open the exact file it was asked to load.

Typical failures look like:

gguf_init_from_file: failed to open GGUF file

or a router response such as:

{"error":{"message":"File Not Found"}}

The shortest useful diagnostic path is:

identify the exact path
-> verify it from the runtime's working directory
-> use an absolute path
-> verify every split shard
-> bypass router/preset indirection

1. Check the path the process is actually using

A path that exists in your shell may not exist relative to the process that launches llama.cpp.

Start with:

pwd
ls -lh ./models
realpath ./models/your-model.gguf
test -r ./models/your-model.gguf && echo readable

If realpath fails or test -r prints nothing, fix that before investigating anything else.

A relative path such as:

models/model.gguf

means:

<current working directory>/models/model.gguf

It does not mean “the models directory I intended.”

2. Use an absolute path as a diagnostic

For a direct load, temporarily remove path ambiguity:

llama-cli -m /absolute/path/to/model.gguf -p "hello"

or:

llama-server -m /absolute/path/to/model.gguf

If the absolute path works while the relative path fails, the model itself was not the problem.

You found a launch-directory or configuration problem.

This is especially useful with service managers, containers, wrappers, launch scripts, or routers because their working directory may differ from an interactive terminal.

3. A split GGUF needs every shard

Large GGUFs may be split into files such as:

model-00001-of-00013.gguf
model-00002-of-00013.gguf
...
model-00013-of-00013.gguf

Having the first shard is not enough.

Check the directory:

find /absolute/path/to/model-dir -maxdepth 1 -type f -name '*.gguf' -printf '%f
' | sort

If llama.cpp reports that a later shard cannot be opened, the first shard was found successfully.

The failure is now narrower:

missing shard
wrong filename
wrong directory
partial download
unreadable shard

Do not solve a missing shard by renaming unrelated files.

The numbered set needs to match what the split metadata expects.

4. Router 404s can hide the path that failed

A router or preset can return a generic:

File Not Found

without making the failing filesystem path obvious.

When that happens, reproduce the same model with the simplest direct command you can.

For example:

llama-server -m /absolute/path/to/model.gguf

If the direct command succeeds, move outward one layer at a time:

direct llama-server
-> preset
-> models directory
-> router
-> client request

This isolates whether the failure belongs to llama.cpp itself or to the model-selection layer around it.

5. Do not assume –models-dir fixes every relative path

Current llama.cpp routing and preset behavior has evolved.

A models directory can help discovery, but it should not be used as proof that every relative -m path is resolved the way you expect.

For troubleshooting, an absolute model path is the clean control.

Once the model loads, you can reintroduce:

--models-dir
presets
aliases
router model names
relative paths

one at a time.

6. Check permissions and ownership

A file can exist but still be unavailable to the account running the service.

Useful checks are:

ls -l /absolute/path/to/model.gguf
namei -l /absolute/path/to/model.gguf
test -r /absolute/path/to/model.gguf && echo readable

The process needs access not only to the file but also to the parent directories in the path.

This matters when the interactive shell runs as one user and the service runs as another.

7. Containers need the file inside the container namespace

A host path such as:

/home/me/models/model.gguf

does not automatically exist inside a container.

If llama.cpp runs in Docker or another container runtime, verify the mounted path from inside that environment.

The important distinction is:

host path
!=
container path

A correct host download does not prove the container can see it.

8. Distinguish File Not Found from model incompatibility

These errors point to different layers.

Missing path

failed to open GGUF file
File Not Found

Start with filesystem and routing.

Unsupported architecture

unknown model architecture

The file was opened far enough for metadata to be read.

Now investigate runtime support.

Missing tensor or wrong tensor shape

tensor ... not found
tensor ... has wrong shape

The model was opened and parsed further still.

Now investigate conversion, metadata, or model/runtime compatibility.

KV-cache allocation failure

failed to allocate buffer for kv cache

The model loaded much further.

Now investigate memory, context length, cache type, and offload.

Treating all four as the same “model won’t load” problem wastes time.

A compact decision tree

File Not Found
|
+-- Does the absolute file path exist?
|   |
|   +-- no -> fix path/download
|   |
|   +-- yes
|       |
|       +-- Is it readable by the runtime user?
|       |   |
|       |   +-- no -> fix permissions/path traversal
|       |
|       +-- Is it a split GGUF?
|       |   |
|       |   +-- yes -> verify every numbered shard
|       |
|       +-- Does direct llama-server load it?
|           |
|           +-- yes -> inspect router/preset/models-dir
|           |
|           +-- no -> use the direct error to continue diagnosis

The fastest practical fix

When someone sends only:

llama.cpp file not found

the first command I want is not a rebuild.

It is:

realpath /path/to/model.gguf
test -r /path/to/model.gguf && echo readable
llama-server -m /absolute/path/to/model.gguf

If the model is split, verify every shard immediately after that.

Only once direct loading succeeds should you spend time on aliases, presets, routing, or automatic model discovery.

That keeps a filesystem error from turning into an unnecessary runtime-debugging session.

Sources and further reading

Continue reading