llama.cpp File Not Found: Fix GGUF Paths, Shards, and Router 404s
Diagnose llama.cpp File Not Found errors by checking the actual GGUF path, process working directory, split-model shards, router configuration, and file readability before blaming model compatibility.
Approximately 5 min read
A llama.cpp File Not Found error is often simpler than it looks.
Before changing quantization, rebuilding CUDA, or blaming the model architecture, first prove that the process can open the exact file it was asked to load.
Typical failures look like:
gguf_init_from_file: failed to open GGUF file
or a router response such as:
{"error":{"message":"File Not Found"}}
The shortest useful diagnostic path is:
identify the exact path
-> verify it from the runtime's working directory
-> use an absolute path
-> verify every split shard
-> bypass router/preset indirection
1. Check the path the process is actually using
A path that exists in your shell may not exist relative to the process that launches llama.cpp.
Start with:
pwd
ls -lh ./models
realpath ./models/your-model.gguf
test -r ./models/your-model.gguf && echo readable
If realpath fails or test -r prints nothing, fix that before investigating anything else.
A relative path such as:
models/model.gguf
means:
<current working directory>/models/model.gguf
It does not mean “the models directory I intended.”
2. Use an absolute path as a diagnostic
For a direct load, temporarily remove path ambiguity:
llama-cli -m /absolute/path/to/model.gguf -p "hello"
or:
llama-server -m /absolute/path/to/model.gguf
If the absolute path works while the relative path fails, the model itself was not the problem.
You found a launch-directory or configuration problem.
This is especially useful with service managers, containers, wrappers, launch scripts, or routers because their working directory may differ from an interactive terminal.
3. A split GGUF needs every shard
Large GGUFs may be split into files such as:
model-00001-of-00013.gguf
model-00002-of-00013.gguf
...
model-00013-of-00013.gguf
Having the first shard is not enough.
Check the directory:
find /absolute/path/to/model-dir -maxdepth 1 -type f -name '*.gguf' -printf '%f
' | sort
If llama.cpp reports that a later shard cannot be opened, the first shard was found successfully.
The failure is now narrower:
missing shard
wrong filename
wrong directory
partial download
unreadable shard
Do not solve a missing shard by renaming unrelated files.
The numbered set needs to match what the split metadata expects.
4. Router 404s can hide the path that failed
A router or preset can return a generic:
File Not Found
without making the failing filesystem path obvious.
When that happens, reproduce the same model with the simplest direct command you can.
For example:
llama-server -m /absolute/path/to/model.gguf
If the direct command succeeds, move outward one layer at a time:
direct llama-server
-> preset
-> models directory
-> router
-> client request
This isolates whether the failure belongs to llama.cpp itself or to the model-selection layer around it.
5. Do not assume –models-dir fixes every relative path
Current llama.cpp routing and preset behavior has evolved.
A models directory can help discovery, but it should not be used as proof that every relative -m path is resolved the way you expect.
For troubleshooting, an absolute model path is the clean control.
Once the model loads, you can reintroduce:
--models-dir
presets
aliases
router model names
relative paths
one at a time.
6. Check permissions and ownership
A file can exist but still be unavailable to the account running the service.
Useful checks are:
ls -l /absolute/path/to/model.gguf
namei -l /absolute/path/to/model.gguf
test -r /absolute/path/to/model.gguf && echo readable
The process needs access not only to the file but also to the parent directories in the path.
This matters when the interactive shell runs as one user and the service runs as another.
7. Containers need the file inside the container namespace
A host path such as:
/home/me/models/model.gguf
does not automatically exist inside a container.
If llama.cpp runs in Docker or another container runtime, verify the mounted path from inside that environment.
The important distinction is:
host path
!=
container path
A correct host download does not prove the container can see it.
8. Distinguish File Not Found from model incompatibility
These errors point to different layers.
Missing path
failed to open GGUF file
File Not Found
Start with filesystem and routing.
Unsupported architecture
unknown model architecture
The file was opened far enough for metadata to be read.
Now investigate runtime support.
Missing tensor or wrong tensor shape
tensor ... not found
tensor ... has wrong shape
The model was opened and parsed further still.
Now investigate conversion, metadata, or model/runtime compatibility.
KV-cache allocation failure
failed to allocate buffer for kv cache
The model loaded much further.
Now investigate memory, context length, cache type, and offload.
Treating all four as the same “model won’t load” problem wastes time.
A compact decision tree
File Not Found
|
+-- Does the absolute file path exist?
| |
| +-- no -> fix path/download
| |
| +-- yes
| |
| +-- Is it readable by the runtime user?
| | |
| | +-- no -> fix permissions/path traversal
| |
| +-- Is it a split GGUF?
| | |
| | +-- yes -> verify every numbered shard
| |
| +-- Does direct llama-server load it?
| |
| +-- yes -> inspect router/preset/models-dir
| |
| +-- no -> use the direct error to continue diagnosis
The fastest practical fix
When someone sends only:
llama.cpp file not found
the first command I want is not a rebuild.
It is:
realpath /path/to/model.gguf
test -r /path/to/model.gguf && echo readable
llama-server -m /absolute/path/to/model.gguf
If the model is split, verify every shard immediately after that.
Only once direct loading succeeds should you spend time on aliases, presets, routing, or automatic model discovery.
That keeps a filesystem error from turning into an unnecessary runtime-debugging session.