How to Run an LLM Locally: A Beginner's Guide
Learn how to run an LLM locally on Windows, Linux, or macOS, choose a model and quantization, understand hardware needs, and avoid common mistakes.
Approximately 21 min read
Running a large language model locally means the model executes on your own computer instead of sending every prompt to a hosted AI service.
You download the model weights, load them with an inference program, and generate responses using your own CPU, GPU, RAM, or unified memory.
For beginners, this can sound much more complicated than it actually is.
You do not need to build an AI model from scratch.
You usually need only three things:
- an inference application or runtime
- a compatible model
- enough memory to run it
This guide explains the entire process, from choosing software and model size to understanding GGUF files, quantization, VRAM, context length, privacy, and common problems.
What does “running an LLM locally” mean?
A hosted AI service generally works like this:
Your computer
↓
Internet
↓
Remote AI server
↓
Model inference
↓
Response
A local LLM changes the architecture:
Your computer
↓
Local inference runtime
↓
Local model weights
↓
Response
Once the required model and runtime files are downloaded, basic inference can happen on your own machine.
Whether a particular application makes additional network requests depends on the application and configuration, so “local model inference” and “the entire application is permanently offline” are not necessarily identical claims.
For example, software may still use the internet to search for models, download updates, or retrieve metadata.
The important distinction is that the language model itself can execute on your hardware.
Why run an LLM locally?
Cloud AI services are convenient, but local inference gives you a different set of trade-offs.
More control over data flow
A local model can process prompts without sending the inference request to a third-party model server.
This can be valuable when working with private notes, code, internal documents, or experiments where you want greater control over where data is processed.
You should still inspect the privacy behavior of the application around the model, especially if you enable plugins, web search, telemetry, remote APIs, or other integrations.
No per-token API bill for local inference
After buying the hardware and downloading the model, local generation does not normally have a per-token cloud inference charge.
It still consumes electricity and hardware resources, so local inference is not literally free.
Offline use
Many local inference tools can generate text without an active internet connection after the necessary application, runtime, and model files are installed.
This is useful for travel, isolated systems, unreliable internet connections, and private environments.
More control over model selection
Instead of using only the models offered by one hosted provider, you can choose among many downloadable open-weight models and quantizations.
You can experiment with different:
model families
parameter sizes
quantizations
context settings
system prompts
inference runtimes
Learning
Local LLMs expose concepts that cloud chat interfaces often hide.
You quickly learn the practical meaning of:
parameters
VRAM
RAM
quantization
GGUF
context length
KV cache
GPU offloading
tokens per second
That makes local AI an excellent way to understand how modern LLM inference actually works.
What hardware do you need?
There is no single minimum specification for all local LLMs.
A tiny model can run on hardware that would be unusable for a much larger model.
The most important resources are:
GPU VRAM
system RAM
memory bandwidth
CPU performance
storage
For most beginners, available memory is the first constraint to check.
You do not need a GPU
A common misconception is:
You need an NVIDIA GPU to run an LLM locally.
You do not.
Software such as llama.cpp can run models on CPUs and supports multiple hardware backends.
A GPU can dramatically improve performance, but CPU inference is completely legitimate.
This means you can start learning local AI even if your computer has no powerful discrete GPU.
The trade-off is usually speed.
Large models running mostly on a CPU may generate tokens much more slowly than the same model running fully on a suitable GPU.
Why GPUs help
LLM inference performs large amounts of matrix computation and moves substantial quantities of model data through memory.
Modern GPUs are designed for highly parallel workloads and typically provide much higher memory bandwidth than ordinary desktop system memory.
If enough model data fits in VRAM, GPU acceleration can make local inference much faster.
For NVIDIA GPUs, many inference systems use CUDA.
Apple Silicon systems commonly use Metal or MLX-compatible runtimes.
AMD and Intel support depends on the operating system, runtime, and backend.
Always check the current runtime documentation before buying hardware solely for a particular inference setup.
How much VRAM do you need?
The answer depends on:
model size
weight precision
quantization
context length
KV cache
runtime
concurrency
For a theoretical dense model, weight storage starts roughly with:
FP16/BF16 ≈ 2 bytes per parameter
8-bit ≈ 1 byte per parameter
4-bit ≈ 0.5 byte per parameter
So an idealized 8-billion-parameter model at 4-bit precision would begin around:
8 billion × 0.5 byte
≈ 4 GB
for weight storage.
Actual runtime memory is higher.
The system also needs memory for the KV cache, temporary buffers, runtime overhead, and other allocations.
For a detailed explanation, read How Much VRAM Do You Need for Local LLMs?.
System RAM matters too
Even when using a GPU, system RAM remains important.
It may be used for:
loading model files
CPU inference
CPU/GPU hybrid inference
other applications
model management
memory-mapped data
If a model does not fully fit in VRAM, some runtimes can keep part of the model in system RAM while accelerating other parts on the GPU.
This is called partial GPU offloading or CPU+GPU hybrid inference.
It allows you to run models larger than your GPU’s VRAM capacity, although performance is usually lower than keeping the entire model on a fast GPU.
Storage requirements
Local models can be large.
Depending on parameter count and quantization, one model may consume several gigabytes or tens of gigabytes of storage.
If you experiment frequently, you may quickly accumulate:
multiple model families
multiple sizes
multiple quantizations
old versions
embedding models
vision models
speech models
An SSD is strongly preferable to slow storage for a pleasant workflow.
You should also leave enough free disk space for temporary downloads and model updates.
Step 1: choose how you want to run the model
There are several good ways to start.
Three common approaches are:
LM Studio
Ollama
llama.cpp
They overlap, but they target somewhat different workflows.
Option 1: LM Studio
LM Studio is one of the easiest choices for someone who wants a graphical interface.
It can search for downloadable models, load them, provide a chat interface, and expose a local server.
Its current desktop application supports Windows, macOS, and Linux, subject to platform-specific hardware and operating-system requirements.
A typical beginner workflow is:
Install LM Studio
↓
Search for a model
↓
Download a compatible version
↓
Load the model
↓
Open Chat
↓
Start prompting
This avoids most command-line work.
LM Studio is particularly useful when you want to experiment with several models without manually managing every command-line option.
Option 2: Ollama
Ollama provides a streamlined model-management and local-serving workflow.
It is available on macOS, Windows, and Linux.
A major attraction is that model setup and execution can be handled through simple commands rather than manually downloading and pointing a runtime at individual model files.
Depending on the current Ollama release, the user interface and command workflow may evolve, so check the official quickstart for the latest commands.
Ollama also exposes a local API, making it useful for connecting local models to applications and development tools.
Conceptually:
Application
↓
localhost API
↓
Ollama
↓
local model
That can be easier than integrating directly with a lower-level inference runtime.
Option 3: llama.cpp
llama.cpp provides more direct control.
It is a C/C++ inference project designed to run LLMs across a wide variety of hardware.
Current llama.cpp supports backends including:
CPU
CUDA
Metal
HIP
Vulkan
SYCL
and others
It can also split inference between CPU and GPU when a model is larger than available VRAM.
llama.cpp is an excellent choice if you want to understand what your local inference stack is actually doing.
It is also widely used underneath higher-level local AI tools.
Which should a beginner choose?
If you want the easiest visual experience, start with:
LM Studio
If you want easy model management plus a local API, consider:
Ollama
If you want direct control and want to learn the underlying inference workflow, start with:
llama.cpp
There is no requirement to choose only one forever.
Many local-AI users eventually keep more than one runtime installed.
Step 2: choose a model
The software is only the inference engine.
You still need a model.
A model name often tells you several useful things.
You might encounter names containing:
1B
3B
7B
8B
14B
32B
70B
The B generally refers to billions of parameters.
An 8B model has roughly eight billion parameters.
Parameter count is not a complete measure of quality, but it strongly affects storage and memory requirements.
Larger models generally require more memory.
Do not start with the largest model you can find
Beginners often assume:
70B must be better than 8B
therefore I need 70B
That is not a good hardware-selection rule.
A smaller modern model can be far easier to run and may be completely adequate for your actual task.
A large model that generates one token every several seconds may be less useful to you than a smaller model that responds quickly.
Start with a model that fits comfortably.
Once the workflow works, move upward.
Base models vs instruction models
You may see terms such as:
Base
Instruct
Chat
IT
A base model is generally the pretrained foundation before instruction tuning.
An instruct or chat model has usually been further trained to follow user instructions and conversational formats.
For a beginner who wants a ChatGPT-like experience, an instruction-tuned model is usually the more appropriate starting point.
Do not accidentally download a base model and then conclude that local LLMs are terrible at following instructions.
Step 3: understand model formats
Models can be distributed in several formats.
Two common things you may encounter are:
GGUF files
Transformers-style model repositories
The correct format depends on your runtime.
For llama.cpp, GGUF is the standard model format.
A model file may look like:
model-Q4_K_M.gguf
Other inference stacks may load models from directories containing files such as:
config.json
tokenizer files
*.safetensors
The model format and inference runtime must be compatible.
Do not download a random model file before deciding which runtime you intend to use.
What is GGUF?
GGUF is a model file format used by llama.cpp and the broader GGML ecosystem.
A GGUF file can package model tensors and metadata in a format that llama.cpp can load.
GGUF models are extremely common in local AI because they can be distributed as convenient files and are available in many quantization levels.
For a deeper comparison with another common deployment option, read GGUF vs AWQ: What’s the Difference and Which Should You Use?.
Step 4: understand quantization
Full-precision model weights consume a large amount of memory.
Quantization reduces the numerical precision used to represent weights so that the model requires less storage and memory.
This can turn a model that would otherwise need tens of gigabytes of memory into something much more practical on consumer hardware.
You may encounter GGUF quantization labels such as:
Q4_K_M
Q5_K_M
Q6_K
Q8_0
These labels represent different quantization configurations.
They are not simply quality scores.
Generally, more aggressive compression reduces memory usage but can also alter model output quality.
The best choice depends on:
available RAM or VRAM
model
runtime
quality requirements
speed
Is Q4_K_M always the best choice?
No.
Q4_K_M is popular because it often provides a useful memory-versus-quality compromise, but there is no universal rule that it is the optimal quantization for every model and workload.
If memory is plentiful, you may prefer a larger quantization.
If memory is extremely constrained, you may need something smaller.
Treat quantization as an engineering trade-off rather than a ranking system.
Step 5: download the model
How you download a model depends on your application.
LM Studio provides model discovery and downloading inside the application.
Ollama manages models through its own workflow.
llama.cpp can load a local GGUF file and can also download compatible models from Hugging Face using its Hugging Face model options.
For a local GGUF file, the basic llama.cpp concept is simple:
llama-cli -m model.gguf
Current llama.cpp can also obtain compatible models from Hugging Face using -hf.
For example, its official documentation demonstrates the pattern:
llama-cli -hf owner/model-GGUF
Always use the exact repository and model variant that you actually intend to run.
Step 6: load the model
Downloading a model and loading a model are different operations.
Downloading means:
model file
→ disk
Loading means making the model available to the inference engine using system memory, GPU memory, or both.
A large model may fit comfortably on disk but still fail to load because there is not enough RAM or VRAM.
For example:
100 GB free SSD space
does not mean:
a 40 GB model fits into a 12 GB GPU
Storage and working memory are separate constraints.
Step 7: choose a reasonable context length
Context length controls how much token history the model can consider within an active request or conversation.
You may see values such as:
4K
8K
16K
32K
128K
Long context is useful, but it is not free.
The model stores attention state in a KV cache during autoregressive inference.
Longer active context generally means more cache memory.
This is why a model can load successfully and later run out of VRAM when you paste a very long document.
For a detailed explanation, read What Is a KV Cache?.
Do not automatically select maximum context
Suppose a model supports 128K tokens.
If your normal conversations use only several thousand tokens, configuring an enormous context window may offer little practical value while increasing memory pressure depending on the runtime and cache implementation.
Choose a context that matches your actual task.
If you later need more, increase it deliberately.
Step 8: test with a simple prompt
Do not begin your first local-model test with an enormous PDF, complex agent system, or benchmark suite.
First verify the basic inference path.
Try something simple:
Explain why the sky appears blue in three short paragraphs.
Check whether:
the model loads
the response is coherent
generation speed is acceptable
memory remains stable
the chat template works
Once the basic path works, increase complexity.
What is a chat template?
Instruction-tuned models are usually trained around a particular conversation structure.
Internally, a chat might need to be represented using special formatting for roles such as:
system
user
assistant
The exact format depends on the model.
Modern model files and runtimes often contain or infer the correct chat template automatically.
But incorrect template handling can cause surprisingly poor behavior.
If a respected instruction model produces strange role markers, ignores instructions, or behaves like text completion rather than chat, template compatibility is one thing to investigate.
Why is my local model slow?
A local LLM can be slow for many different reasons.
The model may be:
too large for your hardware
mostly running on CPU
partially offloaded
using a slow memory path
using an inefficient backend
processing a very long context
running at an unsuitable configuration
Do not judge performance only by parameter count.
Inference speed depends heavily on memory bandwidth, hardware backend, quantization, runtime implementation, prompt processing, and generation settings.
We will cover this separately in the Troubleshooting guide:
Why Is My Local LLM Slow? Common Causes and Fixes
Prompt processing speed and generation speed are different
An LLM workload has multiple phases.
A long prompt first needs to be processed.
After that, the model generates output token by token.
These operations have different performance characteristics.
You may therefore see statistics that distinguish between:
prompt processing / prefill speed
generation / decode speed
A model can process prompts quickly but generate tokens more slowly, or vice versa.
When comparing local inference performance, make sure you know which number is being reported.
What does tokens per second mean?
Generation speed is commonly measured in tokens per second.
A token is not exactly one word.
Depending on the tokenizer and language, one word may contain one token or several tokens.
So:
20 tokens/second
does not simply mean:
20 words/second
Tokens per second is still useful for comparing the same or similar workloads, but benchmark methodology matters.
CPU-only vs full-GPU vs hybrid inference
There are three broad configurations.
CPU-only
Model
→ system RAM
→ CPU computation
Advantages include broad accessibility and the ability to use large amounts of system RAM.
The disadvantage is usually lower performance.
Full GPU
Model
→ GPU VRAM
→ GPU computation
If the model and working inference state fit, this is often the fastest configuration for consumer systems with a capable GPU.
CPU + GPU hybrid
part of model
→ GPU
remaining part
→ CPU/system RAM
This can make models usable even when they are larger than available VRAM.
The cost is that inference can become more dependent on system memory bandwidth, CPU performance, and transfers.
Can you run a 70B model locally?
Yes, under the right configuration.
But “can run” and “runs comfortably” are not equivalent.
A 70B-class dense model is large.
At a theoretical 4 bits per parameter, the weights alone are roughly:
70 billion × 0.5 byte
≈ 35 GB
before additional quantization overhead and inference memory.
A single 24 GB GPU therefore cannot normally keep an ordinary 4-bit 70B model entirely resident in VRAM.
But a runtime supporting CPU+GPU hybrid inference may place part of the model in system RAM.
This makes local execution possible at a lower speed.
What model size should a beginner start with?
A practical strategy is to start with a small-to-medium instruction model that comfortably fits your hardware.
For many modern PCs, a quantized model in approximately the 3B-to-8B range is a low-friction place to begin.
The exact choice depends on:
available memory
task
language
model architecture
quantization
runtime compatibility
Do not interpret that range as a universal hardware requirement or quality recommendation.
Its advantage is simply that smaller models make troubleshooting easier.
Once everything works, try larger models and compare whether the additional resource use produces enough benefit for your workload.
Local model privacy: what it does and does not guarantee
Running inference locally can greatly reduce the need to send prompts to a remote model provider.
But do not assume that the word “local” automatically guarantees complete privacy.
Your overall application may still contain:
analytics
automatic update checks
model search
web search
plugins
remote MCP servers
cloud sync
external APIs
If privacy is important, audit the entire data path.
A model executing on your GPU is one component of that path.
Can a local LLM access the internet?
Not by itself simply because it is an LLM.
A plain local language model receives input and generates output.
For it to search the web or interact with external systems, the surrounding application must provide tools or network capabilities.
Conceptually:
LLM
+
tool framework
+
search/API access
=
internet-enabled assistant
The language model itself does not magically browse the internet.
This distinction becomes important when building local agents.
Can a local LLM read files?
Again, the model itself does not automatically have unrestricted access to your filesystem.
An application can read a file, convert its contents into model input, use retrieval, or expose filesystem tools.
That capability comes from the software around the model.
This is good from a security perspective because permissions can be controlled separately from inference.
Local LLM vs local AI application
It helps to separate three layers:
Model
Runtime
Application
The model contains learned weights.
The runtime executes the model.
The application provides the user interface and additional features.
For example, a simplified architecture might be:
Chat interface
↓
llama.cpp
↓
GGUF model
Or:
Your Python application
↓
Ollama local API
↓
local model
Understanding these layers makes troubleshooting much easier.
If something fails, ask:
Is the model incompatible?
Is the runtime failing?
Is the application misconfigured?
Local LLM vs cloud API
Neither is universally better.
Cloud APIs offer major advantages:
no local GPU requirement
easy scaling
large hosted models
minimal setup
managed infrastructure
Local inference offers different advantages:
control
offline capability
local processing
hardware experimentation
no per-token cloud inference bill
Many practical AI systems use both.
For example:
small routine tasks
→ local model
difficult occasional tasks
→ hosted model
A hybrid strategy can make more sense than trying to force every workload into one environment.
Common beginner mistake: downloading before checking compatibility
Imagine you download a 30 GB model.
Then you discover:
wrong format
wrong runtime
not enough memory
wrong model architecture
unsupported backend
You have wasted time and bandwidth.
Use this order instead:
1. Choose runtime
2. Check hardware support
3. Choose model
4. Choose quantization
5. Check expected memory
6. Download
That simple sequence prevents many beginner problems.
Common beginner mistake: confusing disk size with VRAM
A 12 GB GGUF file does not necessarily require exactly 12 GB of VRAM.
And a GPU with 12 GB of VRAM does not necessarily mean every 12 GB model file will fit entirely on GPU.
Inference needs additional memory.
The model may also be partially placed in system RAM.
Treat the file size as an important starting point, not a complete VRAM calculation.
Common beginner mistake: selecting the biggest context
A model may advertise a huge context window.
That is a capability ceiling, not a recommendation that you use the maximum on every request.
Long contexts can increase memory use and processing cost.
Start smaller.
Increase context when your task actually needs it.
Common beginner mistake: comparing models only by parameter count
Parameter count is useful.
It is not sufficient.
Two models of the same size can differ substantially in:
training data
architecture
instruction tuning
tokenizer
context behavior
language ability
reasoning style
tool support
multimodal support
A newer 8B model may outperform an older or poorly matched larger model on your specific task.
Test the task you actually care about.
Common beginner mistake: trusting random benchmark numbers
Local-AI benchmark results are highly sensitive to configuration.
A useful performance claim should specify details such as:
exact model
quantization
runtime version
CPU
GPU
RAM
context
batch settings
prompt length
generation length
Without those details, a tokens-per-second number may tell you very little about your own machine.
This is why RAMGPT’s Benchmark section should use reproducible measurements rather than fabricated or copied numbers.
A simple beginner workflow
A low-risk first local-LLM experiment looks like this:
Choose LM Studio, Ollama, or llama.cpp
↓
Start with a small instruction model
↓
Choose a moderate quantization
↓
Check model memory requirements
↓
Download
↓
Load with a modest context size
↓
Run simple prompts
↓
Observe RAM/VRAM and speed
↓
Try a larger model only if needed
This is much more reliable than beginning with the largest model your storage drive can hold.
When should you use llama.cpp?
llama.cpp is especially useful when you want:
GGUF models
direct command-line control
CPU inference
GPU acceleration
CPU/GPU hybrid inference
a local server
portable deployment
Its official quickstart can run a local GGUF model with:
llama-cli -m my_model.gguf
It can also launch a server using llama-server.
That makes it useful both for interactive local use and as a backend for other applications.
When should you use LM Studio?
LM Studio makes sense when you want:
graphical model discovery
simple downloading
a desktop chat interface
runtime management
local server capability
minimal command-line work
It is a convenient way to learn what model size and quantization work well on your machine before building a more customized stack.
When should you use Ollama?
Ollama is attractive when you want:
simple model management
local API access
application integration
command-line workflows
It can serve as a convenient boundary between your applications and the underlying local models.
Instead of every project implementing model loading itself, your program can call a localhost endpoint.
Should you use all three?
You can.
There is nothing unusual about using:
LM Studio for interactive testing
llama.cpp for direct experiments
Ollama for application integration
You may eventually decide that one tool covers all your needs.
But beginners benefit from understanding that the model and runtime are separate.
The same general local-AI concepts transfer between tools.
What should you learn next?
Once you can successfully run one model locally, four concepts will unlock most of the rest of the ecosystem.
First, learn quantization.
This explains why the same model can appear in files with dramatically different sizes.
Second, understand VRAM planning.
Read How Much VRAM Do You Need for Local LLMs?.
Third, understand KV cache and context length.
Read What Is a KV Cache? How It Works and Why It Uses VRAM.
Fourth, understand model formats and runtimes.
Read GGUF vs AWQ: What’s the Difference and Which Should You Use?.
Once those pieces make sense, local AI stops looking like a collection of mysterious acronyms and starts looking like a normal software stack.
Bottom line
Running an LLM locally does not require building an AI model.
You need:
a compatible model
+
an inference runtime
+
enough memory
For the easiest graphical start, LM Studio is a practical option.
For simple model management and local API workflows, Ollama is worth considering.
For direct control over GGUF inference, hardware backends, and CPU/GPU offloading, llama.cpp is one of the most important runtimes to understand.
Start with a model that fits comfortably rather than the largest model possible.
Use quantization to control memory requirements.
Keep context length reasonable.
Watch both RAM and VRAM.
And remember that a local AI stack has separate layers:
application
→ runtime
→ model
→ hardware
Once your first model runs successfully, changing models and experimenting with different local AI tools becomes much easier.