Why Is My Local LLM Slow? Common Causes and Fixes
Learn why a local LLM runs slowly and how model size, GPU offloading, context length, memory bandwidth, CPU threads, and runtime settings affect speed.
Approximately 17 min read
Your local LLM works, but it feels painfully slow.
Maybe the model takes a long time before the first word appears. Maybe prompt processing is fast but generation crawls at only a few tokens per second. Or perhaps the GPU is installed but seems barely active.
There is usually a reason.
Local LLM performance depends on much more than the model’s parameter count.
A useful mental model is:
Local LLM speed
=
model
+ quantization
+ hardware
+ memory bandwidth
+ model placement
+ context length
+ runtime
+ workload
The first step is to identify which part is slow.
Do not start changing random settings before separating prompt processing from token generation.
First: identify what “slow” means
LLM inference has two major phases:
Prefill
→ process the prompt
Decode
→ generate output tokens
These can behave very differently.
For example:
Prompt processing: 800 tokens/s
Generation: 25 tokens/s
is not contradictory.
Prompt processing can process many tokens in parallel.
Autoregressive generation produces new tokens sequentially and repeatedly reads model data and attention state.
So before troubleshooting, ask:
Is the delay before generation starts, or is each generated token slow?
Time to first token vs generation speed
A long pause before the first response can be caused by:
- a long prompt
- model loading
- prompt processing
- large context
- cold GPU initialization
- server queueing
- disk loading
- context reuse not being available
Slow text after generation begins points more toward:
- model size
- CPU inference
- incomplete GPU offloading
- memory bandwidth
- runtime configuration
- thermal or power limits
- concurrency
These are different problems.
Check actual performance instead of guessing
If you use llama.cpp, you can enable internal performance timing:
llama-cli \
-m model.gguf \
--perf
Current llama.cpp also includes llama-bench, which is specifically designed for repeatable performance testing.
For example:
llama-bench -m model.gguf
It reports separate performance measurements for prompt processing and text generation.
That is much more useful than saying:
This model feels slow.
Cause 1: the model is too large
The most obvious cause is often the correct one.
A larger model requires more memory movement and computation.
Consider:
8B model
32B model
70B model
Even if all three are quantized, the larger models contain much more weight data.
During generation, those weights must repeatedly participate in computation.
A 70B model should not be expected to generate at the same speed as an 8B model on the same consumer GPU.
If speed matters more than marginal model quality, using a smaller model may be the most effective optimization.
Fix: test a smaller model
Do not change ten runtime settings at once.
Run a smaller model using approximately the same environment.
For example:
current model: 32B Q4
test model: 8B Q4
If generation becomes dramatically faster, the hardware is probably functioning correctly.
The larger model is simply a much heavier workload.
For beginners, read How to Run an LLM Locally: A Beginner’s Guide.
Cause 2: the model is running on CPU instead of GPU
A common problem is assuming that the presence of a GPU means the model is using it.
That is not necessarily true.
Your runtime needs:
a supported GPU
+
the correct backend
+
a compatible build
+
correct model placement
If llama.cpp was built without CUDA support on an NVIDIA system, changing GPU-layer settings will not magically enable CUDA.
Check whether llama.cpp sees your GPU
Run:
llama-cli --list-devices
On a properly configured GPU build, the relevant accelerator should appear.
If you built llama.cpp yourself for NVIDIA CUDA, the build configuration normally includes:
-DGGML_CUDA=ON
For example:
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
Then:
./build/bin/llama-cli --list-devices
Do this before tuning anything else.
Cause 3: only part of the model is on the GPU
A model can run with some layers on the GPU and the rest using system memory and CPU resources.
This is one of llama.cpp’s most useful features because it allows models larger than available VRAM to run.
But it can also reduce performance substantially.
Conceptually:
Full GPU:
Model
↓
GPU VRAM
↓
GPU
Hybrid:
Part of model → GPU VRAM
Part of model → system RAM / CPU
Hybrid inference increases dependence on the CPU and system-memory path.
Check GPU-layer placement
Current llama.cpp exposes:
-ngl
--gpu-layers
--n-gpu-layers
You can start with automatic placement:
llama-cli \
-m model.gguf \
-ngl auto
If the model comfortably fits in VRAM, you can test:
llama-cli \
-m model.gguf \
-ngl all
Monitor VRAM carefully.
Do not force full GPU placement if it causes an out-of-memory error.
For installation and GPU setup details, see How to Run a GGUF Model with llama.cpp.
More GPU offload often helps, but test it
More layers on the GPU often improve performance when the GPU is substantially faster than the CPU path.
But the result depends on the hardware.
llama.cpp includes benchmarking support specifically for comparing different GPU-layer counts.
That is better than copying somebody else’s -ngl number.
Their model, GPU, CPU, RAM, and llama.cpp version may be completely different from yours.
Cause 4: your model barely fits in VRAM
A configuration can technically work while still being poorly sized.
Suppose:
GPU VRAM: 24 GB
Working usage: 23.8 GB
You have almost no headroom.
Longer contexts, temporary allocations, additional applications, or runtime changes can push the process over the limit.
Some systems may respond by falling back to different memory paths, failing allocations, or becoming unstable.
A better target is not:
Use every possible byte.
It is:
Leave enough memory for the actual workload.
For detailed sizing, read How Much VRAM Do You Need for Local LLMs?.
Cause 5: the quantization is larger than you need
GGUF models may be offered in variants such as:
Q4_K_M
Q5_K_M
Q6_K
Q8_0
A larger quantization generally stores more weight information and requires more memory.
That can be desirable for model quality.
But if the larger quantization no longer fits well in VRAM, the performance trade-off can become significant.
For example, you may find that:
Q8_0
→ partly CPU/offloaded
Q4_K_M
→ fully GPU resident
In that scenario, the smaller quantization may be much faster because model placement changed.
The important point is that quantization affects more than disk size.
It can determine whether the workload fits efficiently on your hardware.
For background, read GGUF vs AWQ: What’s the Difference and Which Should You Use?.
Cause 6: context length is too large
Context affects both memory and computation.
A conversation containing:
2,000 tokens
is a very different workload from one containing:
80,000 tokens
The runtime maintains KV-cache state for previously processed tokens.
That cache consumes memory and can influence inference performance.
A huge context window is not free simply because the model supports it.
Fix: test with a short context
If you configured:
-c 131072
try a much smaller value for troubleshooting:
-c 8192
For example:
llama-cli \
-m model.gguf \
-c 8192 \
-ngl auto
If performance and memory behavior improve substantially, context configuration was part of the problem.
Read What Is a KV Cache? How It Works and Why It Uses VRAM for the underlying mechanism.
Cause 7: the current conversation itself is very long
Configured maximum context and actual context are related but different.
Suppose your runtime supports 32K tokens.
A brand-new conversation might contain:
500 tokens
while an old session might contain:
28,000 tokens
The old conversation carries much more inference state.
So when performance suddenly becomes worse after chatting for a long time, test a fresh conversation.
If a new session is much faster, the hardware did not necessarily become slower.
The workload became larger.
Cause 8: prompt processing is slow, but generation is fine
This happens especially with long documents.
For example:
20,000-token document
→ long prefill
then
short generated answer
→ acceptable decode speed
Do not diagnose the entire system based only on the initial delay.
Measure prompt processing and generation separately.
llama-bench explicitly distinguishes:
pp = prompt processing
tg = text generation
That separation is useful because optimizations may affect the two phases differently.
Cause 9: CPU threads are poorly configured
If some or all inference occurs on the CPU, thread configuration matters.
Current llama.cpp exposes:
-t
--threads
for generation and separate batch-thread controls.
More threads do not automatically mean more speed.
After a point, additional threads can create:
contention
synchronization overhead
memory-bandwidth saturation
This is why using every logical CPU thread is not always optimal.
Test CPU thread counts
Instead of guessing, use llama-bench.
For example:
llama-bench \
-m model.gguf \
-t 4,8,12,16
Compare results.
The optimal number depends on your CPU architecture, physical cores, memory subsystem, and workload.
Physical cores vs logical threads
SMT or Hyper-Threading creates multiple logical threads per physical core.
That can help some workloads.
But LLM generation is often heavily constrained by memory movement, so doubling logical threads does not necessarily double performance.
If CPU inference appears unexpectedly slow, benchmark several thread counts rather than automatically selecting the highest possible number.
Cause 10: memory bandwidth is the bottleneck
LLM token generation frequently moves a large amount of model data for each generated token.
This makes memory bandwidth extremely important.
That applies to:
GPU VRAM bandwidth
system RAM bandwidth
unified memory bandwidth
A system can have a fast processor but still be limited by how quickly model data moves through memory.
This is one reason two GPUs with similar compute specifications can perform differently on LLM inference.
It is also why CPU inference does not scale indefinitely with more cores.
The memory subsystem eventually becomes the limiting factor.
Why model size matters so much for decode
Consider a highly simplified view:
Generate one token
↓
read/process large portions of model weights
↓
generate next token
↓
repeat
If the model is much larger, more data has to participate in every decoding step.
Quantization helps reduce the amount of weight data that needs to be stored and moved.
That is one reason quantization can improve the practicality of local inference.
Cause 11: another process is using the GPU
Your LLM is not necessarily alone.
GPU memory and compute may also be consumed by:
desktop compositor
browser
game
Stable Diffusion
video generation
CUDA jobs
another LLM
training job
On NVIDIA systems, check:
nvidia-smi
For live monitoring:
watch -n 1 nvidia-smi
Look at:
VRAM usage
GPU utilization
other processes
temperature
power
If another application is using several gigabytes of VRAM, close it and retest.
Cause 12: thermal throttling
A GPU or CPU may start fast and become slower after sustained load.
Possible signs include:
good performance initially
↓
temperature rises
↓
clock speed falls
↓
tokens/s decreases
Monitor temperature and clocks while the model is running.
Problems can result from:
dust
poor airflow
weak cooling
laptop thermal limits
high ambient temperature
fan configuration
If performance deteriorates over time rather than being slow immediately, thermals are worth checking.
Cause 13: power limits
Hardware can also run below its normal performance level because of power policy.
Examples include:
laptop battery mode
reduced GPU power limit
energy-saving profile
quiet mode
data-center power cap
A GPU model name alone does not guarantee a particular inference speed.
The actual clock rates and power state matter.
This is especially important when comparing a laptop GPU with a desktop GPU carrying a similar product name.
Cause 14: your laptop is on battery
For laptops, always test serious LLM performance while connected to appropriate external power.
Battery mode may reduce CPU and GPU power dramatically.
If your laptop produces:
8 tokens/s on battery
and:
25 tokens/s plugged in
the model did not change.
The hardware power budget did.
Cause 15: the wrong backend is being used
Inference runtimes can support multiple acceleration backends.
For llama.cpp, examples include:
CUDA
Metal
Vulkan
HIP
SYCL
CPU
Backend maturity and performance differ across hardware and workloads.
If you expected CUDA but accidentally built CPU-only llama.cpp, no amount of prompt tuning will solve the problem.
Start with:
llama-cli --list-devices
Then inspect the startup log.
Confirm that the backend you intended to use is actually active.
Cause 16: Flash Attention configuration
Current llama.cpp includes:
-fa
--flash-attn
with:
on
off
auto
and currently defaults to automatic behavior.
For compatible models and hardware, Flash Attention can affect memory use and performance.
Do not blindly force it on or off.
A good troubleshooting approach is:
-fa auto
first.
Then benchmark:
-fa on
and:
-fa off
if you are investigating a performance difference.
Actual improvement depends on the model, backend, hardware, and context.
Cause 17: KV cache configuration
The KV cache has its own data type and memory behavior.
Current llama.cpp exposes separate controls for key and value cache types.
For example:
--cache-type-k
--cache-type-v
The default cache representation and the model’s weight quantization are separate concepts.
A model might use:
Q4_K_M weights
while the KV cache uses:
FP16
Changing KV-cache precision can reduce memory consumption when supported.
But it should be tested rather than assumed to make everything faster.
Lower memory use and higher tokens per second are not the same objective.
Cause 18: swapping or memory pressure
If system RAM is nearly exhausted, the operating system may begin reclaiming memory aggressively or using swap.
That can make local inference painfully slow.
Check Linux memory usage:
free -h
and:
vmstat 1
If your system is actively swapping while generating text, reducing memory use can help dramatically.
Options include:
smaller model
smaller quantization
more GPU offload
shorter context
closing other programs
adding RAM
Cause 19: slow storage
Storage matters most when loading models, not normally for steady-state token generation when the necessary working data is already in memory.
If:
model takes 90 seconds to load
but then:
generation is fast
your storage path may be the main startup bottleneck.
A fast SSD improves model loading and general usability.
But replacing an SSD will not necessarily improve token-generation speed if the model already resides in RAM or VRAM during inference.
Separate:
model loading speed
from:
generation speed
before spending money.
Cause 20: too many simultaneous requests
A single-user local chat and an API server are different workloads.
Suppose your server handles:
1 active request
and then:
10 active requests
The model weights can be shared, but active sequences consume additional resources, including KV-cache capacity.
Server scheduling also changes latency.
Higher throughput does not necessarily mean each individual user receives tokens faster.
This distinction is important:
throughput
≠
single-request latency
A serving configuration optimized for many users may behave differently from a desktop chat configuration optimized for one person.
Cause 21: batch settings
Current llama.cpp exposes controls such as:
-b
--batch-size
-ub
--ubatch-size
These particularly influence how work is divided during prompt processing.
Larger is not automatically better.
The useful range depends on available memory, model, backend, and hardware.
If your goal is optimization, benchmark multiple values instead of assuming the maximum is best.
llama-bench supports comparing batch sizes directly.
Cause 22: your expectations come from a different workload
Online benchmark numbers are easy to misinterpret.
Suppose someone reports:
100 tokens/s
That number may represent:
different model
smaller model
different quantization
different GPU
different context
prompt processing rather than generation
multiple GPUs
different runtime
different software version
Comparing that number with your own output may be meaningless.
A valid comparison should specify the environment.
Record the complete benchmark configuration
At minimum, record:
model
model revision
quantization
runtime and version
CPU
GPU
VRAM
RAM
backend
context length
GPU layers
batch settings
prompt size
generated token count
Then compare.
Without those details, performance claims are difficult to reproduce.
Use llama-bench for repeatable tests
If you use llama.cpp, llama-bench is specifically intended for performance testing.
A simple test:
llama-bench -m model.gguf
You can also compare GPU-layer placement:
llama-bench \
-m model.gguf \
-ngl 0,10,20,30
Or CPU threads:
llama-bench \
-m model.gguf \
-t 4,8,12,16
Or prompt sizes:
llama-bench \
-m model.gguf \
-p 512,2048,8192
Change one dimension at a time whenever possible.
A good troubleshooting baseline
Start with a controlled configuration.
For example:
llama-cli \
-m model.gguf \
-c 8192 \
-ngl auto \
-fa auto \
--perf
Then check:
llama-cli --list-devices
and, for NVIDIA:
nvidia-smi
Ask:
Is the GPU detected?
Is the model mostly on GPU?
How much VRAM is used?
How much RAM is used?
Is the system swapping?
What is prompt-processing speed?
What is generation speed?
That gives you a real diagnosis instead of random tuning.
A practical troubleshooting order
Use this order:
1. Confirm the model size
Is the model simply much larger than your hardware can run efficiently?
2. Confirm the backend
Run:
llama-cli --list-devices
Make sure the expected accelerator exists.
3. Check model placement
Confirm whether the model is fully on GPU, partially offloaded, or CPU-only.
4. Check VRAM and RAM
Use:
nvidia-smi
free -h
Look for memory exhaustion.
5. Start a fresh, short conversation
Eliminate long-context effects.
6. Use a moderate context
For example:
-c 8192
rather than immediately using the model’s maximum.
7. Measure prompt and generation separately
Use llama.cpp timing output or llama-bench.
8. Test a smaller model
This quickly reveals whether the workload is simply too large.
9. Test another quantization
Especially if the current model does not fit fully in fast memory.
10. Change one setting at a time
Otherwise you will not know which change actually helped.
Example: 32B model is slow on a 24 GB GPU
Imagine:
GPU: 24 GB
Model: 32B
Quantization: large GGUF quantization
Context: 32K
The model may not fit comfortably in VRAM.
Part may remain in system RAM.
A better troubleshooting experiment could be:
same 32B model
smaller quantization
8K context
automatic GPU placement
fresh conversation
If generation becomes much faster, the problem was not necessarily a bad GPU.
The original workload simply exceeded the efficient full-GPU configuration.
Example: GPU utilization looks low
Low GPU utilization does not automatically mean the GPU is broken.
Possible explanations include:
part of model on CPU
memory-bandwidth limitation
small sequential decode workload
CPU bottleneck
data transfer
synchronization
another part of the pipeline waiting
Do not judge LLM performance solely by the percentage shown in a monitoring utility.
Measure tokens per second and inspect model placement as well.
Example: first response is slow, later responses are fast
Possible reasons include:
initial model loading
GPU initialization
large initial prompt
cold caches
prompt prefill
If subsequent short prompts are fast, generation itself may not be the problem.
Measure model-loading time, time to first token, and decode speed separately.
Example: performance gets worse during a long chat
The active context may be growing.
As the conversation retains more tokens:
KV cache grows
attention workload changes
memory pressure increases
Start a new conversation.
If speed improves, long-context workload was likely contributing.
Do you need a faster GPU?
Maybe, but do not make that the first conclusion.
A GPU upgrade makes sense when you have already established that:
runtime is configured correctly
GPU backend works
model placement is understood
context is reasonable
quantization is appropriate
current hardware is genuinely limiting the workload
Otherwise you may spend money and reproduce the same configuration problem on a faster card.
More VRAM vs faster GPU
These solve related but different problems.
More VRAM helps you:
fit larger models
use larger quantizations
retain longer contexts
avoid CPU offloading
support more concurrent cache state
Higher memory bandwidth and compute capability can improve speed.
A GPU with more VRAM is not automatically faster than every GPU with less VRAM.
When buying hardware for local AI, consider:
VRAM capacity
memory bandwidth
backend compatibility
compute capability
power
cooling
price
not only the VRAM number.
The fastest optimization may be a smaller model
It is tempting to spend hours tuning a 70B model because it technically runs.
But if your task can be solved well by an 8B or 14B model, switching models can produce a much larger speed improvement than fine-tuning a dozen runtime flags.
Local AI is an engineering trade-off.
The objective is not:
Run the largest model possible.
The objective is:
Run a model that provides enough quality at an acceptable speed and memory cost.
Bottom line
When a local LLM is slow, check these first:
1. Is the model too large?
2. Is the GPU actually being used?
3. Is the model fully or partially offloaded?
4. Does the model fit comfortably in VRAM?
5. Is the context unnecessarily large?
6. Is the current conversation very long?
7. Is system RAM exhausted or swapping?
8. Are CPU threads configured reasonably?
9. Is another process using the GPU?
10. Are thermals or power limits reducing performance?
11. Are you comparing prompt speed with generation speed?
12. Are your benchmark settings actually comparable?
For llama.cpp, a useful diagnostic starting point is:
llama-cli \
-m model.gguf \
-c 8192 \
-ngl auto \
-fa auto \
--perf
Then verify devices:
llama-cli --list-devices
and benchmark the model:
llama-bench -m model.gguf
Do not optimize from intuition alone.
Measure the workload, find the bottleneck, change one variable, and measure again.