What Is LLM Quantization? FP16, INT8 and INT4 Explained
Learn how LLM quantization reduces model memory, what FP16, INT8 and INT4 mean, and how GGUF, AWQ, GPTQ, scales, groups, and quality trade-offs work.
Approximately 25 min read
Download a local LLM and you will quickly encounter labels such as:
FP16
BF16
INT8
INT4
Q4_K_M
AWQ
GPTQ
All of them are related to how model numbers are represented or compressed, but they do not all mean the same thing.
The central idea is quantization.
Quantization reduces the numerical precision used to represent some parts of a neural network, usually its weights.
That can dramatically reduce:
model file size
RAM usage
VRAM usage
memory bandwidth requirements
and may also improve inference performance on hardware and runtimes optimized for the quantized format.
But quantization is not free.
Reducing precision can introduce approximation error.
The practical goal is therefore not:
use the fewest bits possible
It is:
use the lowest precision that still gives
acceptable quality and good hardware performance
This guide explains what that means.
What exactly is being quantized?
A large language model contains billions of learned parameters.
These parameters are numerical values called:
weights
A tiny example might contain values such as:
0.13742
-0.83291
0.00417
1.24931
Real LLMs contain billions of numbers organized into large tensors.
During training, these values are normally represented using floating-point formats.
For deployment, storing every weight at high precision can consume enormous amounts of memory.
Quantization replaces or approximates those high-precision numbers using a smaller numerical representation.
Conceptually:
high-precision weights
↓
quantization
↓
lower-precision representation
The model becomes smaller.
Why does precision consume memory?
Every model parameter occupies some number of bits.
For a simple theoretical dense model:
FP32
= 32 bits
= 4 bytes per parameter
FP16
= 16 bits
= 2 bytes per parameter
INT8
= 8 bits
= 1 byte per parameter
INT4
= 4 bits
= 0.5 byte per parameter
So the theoretical weight memory is approximately:
parameters × bits per parameter ÷ 8
For an 8-billion-parameter model:
FP16:
8 billion × 16 / 8
≈ 16 GB
At an idealized 8 bits:
8 billion × 8 / 8
≈ 8 GB
At an idealized 4 bits:
8 billion × 4 / 8
≈ 4 GB
This simple calculation explains why 4-bit models are so attractive for local AI.
But real quantized files are usually somewhat more complicated than this ideal calculation.
Why isn’t a 4-bit model exactly 0.5 byte per parameter?
Quantization usually requires metadata.
A group of weights may need additional values such as:
scale
zero point
group metadata
higher-precision tensors
alignment overhead
Some tensors may intentionally remain at higher precision.
Therefore:
4-bit quantization
does not necessarily mean:
exactly 4.000 bits for every parameter in the file
A GGUF quantization, for example, may have an effective bits-per-weight value greater than the nominal number in its label.
That is normal.
FP32
FP32 means:
32-bit floating point
It has historically been a common numerical format for neural-network training and general computation.
Each value uses four bytes.
For large models, this becomes expensive.
A 70B-parameter model stored entirely as FP32 weights would theoretically require roughly:
70 billion × 4 bytes
=
280 GB
for weights alone.
That does not include runtime memory.
For inference, using FP32 weights for a large modern LLM is therefore often unnecessary and impractical.
FP16
FP16 means:
16-bit floating point
It stores each value in two bytes.
Compared with FP32:
FP32
4 bytes
FP16
2 bytes
so idealized weight storage is cut approximately in half.
For a 70B dense model:
70 billion × 2 bytes
≈ 140 GB
again before runtime overhead.
FP16 preserves a floating-point representation but has less numerical precision and range than FP32.
BF16
BF16 means:
bfloat16
Like FP16, it occupies:
16 bits
per value.
But BF16 and FP16 allocate those bits differently between exponent and significand.
BF16 retains an exponent range similar to FP32 while providing less mantissa precision.
This makes BF16 useful in many modern training and inference workloads where supported by the hardware.
From a simple model-storage perspective:
FP16
≈ 2 bytes/parameter
BF16
≈ 2 bytes/parameter
Neither one should normally be described as aggressive low-bit LLM quantization.
They are reduced-precision floating-point formats.
Is FP16 quantization?
The terminology can become fuzzy.
Going from FP32 to FP16 reduces numerical precision, so in a broad mathematical sense it is a precision reduction.
But when people discuss:
LLM quantization
they usually mean more aggressive compression such as:
8-bit
4-bit
3-bit
2-bit
or specialized floating-point formats.
So in practical local-LLM discussion, FP16 and BF16 are usually treated as higher-precision baselines rather than what users mean when they say:
a quantized model
INT8
INT8 uses:
8-bit integer values
instead of storing every weight directly as a 16- or 32-bit floating-point number.
A naive approach might seem simple:
0.134 → 3
0.295 → 7
0.881 → 21
But converting arbitrary floating-point weights to integers requires a mapping.
That mapping is where:
scales
zero points
ranges
groups
become important.
Quantization maps continuous values to fewer levels
Suppose a group of weights lies between:
-1.0 and +1.0
A high-precision format can represent many values within that interval.
An 8-bit integer representation contains only a limited number of integer codes.
The quantization system therefore establishes a mapping between:
real-valued weights
and:
quantized integer codes
Conceptually:
real weight
↓
scale / mapping
↓
integer code
During inference, computation can use the quantized representation directly or convert information into an appropriate compute representation depending on the kernel.
Scale factor
A common quantization model uses a scale.
In a simplified symmetric example:
quantized value ≈ round(real value / scale)
and approximate reconstruction becomes:
real value ≈ quantized value × scale
The scale determines the spacing between representable values.
This is an approximation.
The original floating-point number may not be exactly representable after quantization.
That difference is called:
quantization error
A tiny example
Suppose:
scale = 0.1
and the original weight is:
0.37
Then:
0.37 / 0.1
=
3.7
Round:
4
Reconstruct:
4 × 0.1
=
0.4
The original value was:
0.37
The approximated value becomes:
0.40
The difference:
0.03
is quantization error.
Real LLM quantization algorithms are far more sophisticated than this small example, but the central idea is the same.
Why fewer bits create more error
Imagine representing values with many possible numerical levels.
High precision gives the model many choices.
Lower precision gives fewer.
Conceptually:
High precision:
|.|.|.|.|.|.|.|.|.|.|.|.|.|.|.|.|
Low precision:
|----|----|----|----|
A weight has to be approximated by one of the available levels.
As precision decreases, the distance between possible representations can become larger.
This increases the risk of changing model behavior.
The challenge is deciding how to quantize weights while minimizing meaningful error.
INT4
INT4 uses approximately four bits to represent quantized values.
Four bits allow:
2^4
=
16
possible bit patterns.
That is dramatically fewer representable states than an 8-bit value.
But the storage advantage is substantial.
Theoretical weight memory becomes approximately one-quarter of a 16-bit baseline:
FP16
16 bits
INT4
4 bits
This is why 4-bit quantization is extremely popular for consumer local LLM inference.
4-bit does not automatically mean terrible quality
A neural network contains many weights, and they do not all contribute equally to output error.
Modern quantization methods take advantage of model structure and weight distributions.
Instead of simply rounding every weight independently using one global scale, techniques can use:
group-wise quantization
per-channel scales
importance information
calibration data
mixed tensor types
error compensation
This can preserve much more model quality than naive rounding.
That is why practical 4-bit LLMs can remain highly useful.
What is group-wise quantization?
Rather than use one scale for an entire enormous tensor, weights can be divided into smaller groups.
For example:
weights 1–128
→ scale A
weights 129–256
→ scale B
weights 257–384
→ scale C
Each group gets a scale that better matches the range of weights inside it.
This can reduce quantization error.
The cost is additional metadata.
Smaller groups generally require more scale values.
So there is a trade-off:
smaller groups
→ potentially better representation
→ more metadata
Per-tensor quantization
The simplest mapping might use one scale for an entire tensor:
Tensor
→ one scale
This has low metadata overhead.
But one scale may poorly represent a tensor whose numerical ranges vary substantially between channels or regions.
Per-channel quantization
A finer strategy can assign separate quantization parameters to channels.
Conceptually:
Channel 1 → scale 1
Channel 2 → scale 2
Channel 3 → scale 3
This gives the quantizer greater flexibility.
Again, the price is additional metadata and implementation complexity.
Symmetric quantization
In symmetric quantization, values are mapped around zero using a symmetric range.
Conceptually:
-negative
↓
zero
↓
positive
A simplified formula may use:
q = round(x / scale)
with zero naturally corresponding to zero.
This can simplify implementation.
Asymmetric quantization
An asymmetric scheme may use both:
scale
+
zero point
This allows the representable integer range to map to a real-valued range that is not symmetric around zero.
Conceptually:
real value
≈
scale × (quantized value - zero point)
Which strategy is better depends on the data distribution, method, and hardware implementation.
Weight-only quantization
An important distinction is:
weight-only quantization
Here, model weights are stored at low precision while activations or intermediate computation may remain at higher precision.
Conceptually:
Weights
→ INT4
Activations
→ FP16/BF16
AWQ is an example of a weight-only quantization method.
This distinction matters because saying:
4-bit model
does not mean every number used during inference is four bits.
What are activations?
Weights are the learned model parameters.
Activations are intermediate values produced when input passes through the network.
Conceptually:
Input
↓
Layer weights
↓
Activations
↓
Next layer
↓
More activations
During inference, activations are created dynamically.
Weight quantization and activation quantization are therefore separate engineering decisions.
Weight-and-activation quantization
Some methods quantize both:
weights
and:
activations
You may see notation such as:
W8A8
meaning approximately:
8-bit weights
8-bit activations
or:
W4A16
meaning:
4-bit weights
16-bit activations
Exact meaning depends on the framework.
These labels provide more information than simply saying:
4-bit
KV cache quantization is another separate thing
An LLM also maintains the KV cache during autoregressive inference.
That cache can use its own numerical type.
For example, you could theoretically have:
model weights:
4-bit quantized
KV cache:
FP16
or a runtime that supports:
model weights:
4-bit
KV cache:
lower precision
These are independent choices.
Read What Is a KV Cache? How It Works and Why It Uses VRAM for a detailed explanation.
Post-training quantization
A very important category is:
Post-Training Quantization
PTQ
The model is first trained at a higher precision.
Then quantization is applied afterward.
Conceptually:
Train model
↓
High-precision model
↓
Quantize
↓
Deployment model
This is attractive because the extremely expensive original training process does not need to be repeated from scratch just to create a quantized deployment version.
AWQ and GPTQ are examples of post-training quantization methods.
Quantization-aware training
Another approach is:
Quantization-Aware Training
QAT
Here, training or fine-tuning simulates or incorporates quantization effects.
The model can adapt to low-precision behavior during optimization.
Conceptually:
training
+
simulated quantization effects
↓
model adapted for low precision
QAT can produce excellent low-precision results, but it requires additional training computation.
For ordinary users downloading local LLMs, post-training quantized models are much more commonly encountered.
What is AWQ?
AWQ stands for:
Activation-aware Weight Quantization
It is a low-bit weight-only quantization method designed for LLM deployment.
A key idea from AWQ is that model weights are not equally important.
The method uses activation statistics to identify important weight channels and reduce harmful quantization error.
Conceptually:
Calibration activations
↓
identify important channels
↓
adjust scaling
↓
quantize weights
AWQ does not mean:
a model file format
It describes a quantization method.
This distinction is important.
For more detail, read GGUF vs AWQ: What’s the Difference and Which Should You Use?.
What is GPTQ?
GPTQ is another post-training weight quantization method for large transformer models.
It uses approximate second-order information to quantize weights while compensating for quantization error.
The practical outcome is that models can be compressed to low bit widths such as:
4-bit
3-bit
while attempting to preserve accuracy.
GPTQ and AWQ therefore solve similar deployment problems using different quantization strategies.
They should not be treated as simple synonyms.
GGUF vs AWQ vs GPTQ
These terms are frequently mixed together.
A useful mental model is:
GGUF
→ model file/container format used heavily in llama.cpp ecosystem
AWQ
→ quantization method
GPTQ
→ quantization method
A comparison like:
GGUF vs AWQ
is therefore somewhat asymmetric.
In practice, users make that comparison because different runtime ecosystems distribute models using these labels.
But conceptually they exist at different layers.
What are GGUF Q4 and Q5 models?
In the llama.cpp ecosystem, GGUF files can use numerous tensor quantization types.
You may encounter filenames such as:
Q4_0
Q4_K_M
Q5_K_M
Q6_K
Q8_0
These are llama.cpp/GGML-family quantization configurations.
The initial number is useful, but it does not completely describe the effective storage or quality.
For example:
Q4_K_M
does not mean:
every tensor uses exactly four bits
The format can use mixtures of tensor types and additional block metadata.
We will cover these variants in a separate guide:
GGUF Quantization Levels Explained:
Q4, Q5, Q6 and Q8
Why Q4_K_M is popular
Q4_K_M is commonly used because it often provides a practical balance among:
model size
memory use
quality
speed
in llama.cpp-style deployments.
But it is not a universal best quantization.
Depending on model, hardware, and quality needs, you may prefer:
Q5
Q6
Q8
or something smaller than Q4.
The correct choice depends on the workload.
Quantization changes model size
Suppose the same model is distributed as:
BF16
Q8
Q6
Q5
Q4
You should expect lower-bit variants to generally require less storage.
This can make previously impossible models practical.
For example:
32B model in BF16
may be far too large for a consumer GPU.
A 4-bit quantized version may fit much more comfortably.
For detailed memory calculations, read How Much VRAM Do You Need for Local LLMs?.
Quantization can change whether a model fits entirely in VRAM
This may be more important than the direct numerical speed difference between quantization formats.
Imagine:
GPU:
24 GB VRAM
Model A:
large quantization
→ does not fit fully
→ CPU/GPU hybrid
Model B:
smaller quantization
→ fits fully in VRAM
→ full GPU inference
Model B can have a major practical performance advantage because the entire placement strategy changes.
Quantization therefore affects:
storage
memory
hardware placement
and potentially speed
Does quantization always make inference faster?
No.
Smaller models reduce memory traffic, which can help performance.
But actual speed also depends on:
hardware
runtime
quantized kernels
model architecture
batch size
context
GPU backend
CPU backend
dequantization overhead
If your hardware has highly optimized kernels for one quantization format, it may perform very well.
Another format might save memory without providing the same speed advantage.
Therefore:
fewer bits
≠
automatically more tokens per second
Benchmark your actual runtime.
Quantized storage vs compute precision
A model may be stored using one precision while computation uses another representation internally.
For example:
stored weights:
4-bit
computation:
FP16/BF16 or specialized integer/mixed kernels
This is implementation-dependent.
Do not assume that a model labeled:
INT4
means the GPU performs every operation using ordinary 4-bit integer arithmetic.
The runtime and kernels decide how data is loaded, unpacked, dequantized, and multiplied.
Why hardware compatibility matters
A quantization format is useful only if your inference runtime and hardware support it efficiently.
For example, a GPU-serving stack may have optimized kernels for specific methods.
A CPU-oriented runtime may prefer a different family of quantization formats.
That is why choosing a quantized model should start with:
runtime
+
hardware
not merely the smallest file available.
Quantization is not compression like ZIP
A ZIP file becomes compressed on disk and is decompressed back into the exact original bytes.
Lossless compression means:
original data
→ compression
→ decompression
→ exact original data
Low-bit LLM quantization is generally different.
It usually changes the numerical representation itself.
Conceptually:
0.37291
→
approximate lower-precision value
The original weight may not be recoverable exactly.
That makes ordinary low-bit quantization a form of lossy numerical compression.
Why can the model still work after billions of values change?
Neural networks have a degree of numerical tolerance and redundancy.
The model’s behavior depends on the collective operation of billions of parameters.
A small approximation in an individual weight does not necessarily cause a meaningful output error.
The real challenge is ensuring that quantization errors do not accumulate in important parts of the network.
Modern methods therefore pay special attention to:
weight distribution
activation behavior
sensitive channels
grouping
scaling
error compensation
Quantization quality is model-dependent
There is no rule saying:
every 4-bit model loses exactly X% quality
The impact depends on:
architecture
parameter count
quantization method
calibration
tensor types
task
evaluation metric
Some models tolerate aggressive quantization better than others.
Some tasks may reveal degradation more clearly.
Therefore, avoid universal quality percentages unless they come from a specific reproducible benchmark.
Larger models can sometimes tolerate quantization surprisingly well
A large model may contain substantial redundancy and capacity.
This means a larger quantized model can sometimes be more useful than a smaller high-precision model within the same memory budget.
For example, the practical choice may be between:
small model at high precision
and:
larger model at lower precision
Which one performs better depends on the models and task.
Parameter count and bit width must be evaluated together.
Calibration data
Some quantization methods use a calibration dataset.
The quantizer feeds representative inputs through the model and observes information such as:
activation ranges
sensitive channels
output error
This helps determine better quantization parameters.
Conceptually:
sample prompts
↓
model activations
↓
quantization statistics
↓
better scaling / quantization decisions
Calibration does not mean retraining the original model from scratch.
Why calibration data matters
If a quantization algorithm uses activation statistics, its decisions depend partly on what it observes.
A calibration dataset should therefore be reasonably representative of expected model behavior.
Poor calibration data can make quantization choices less effective.
How sensitive a method is to calibration data depends on the method.
What is an importance matrix in llama.cpp?
The llama.cpp quantization tooling can optionally use an:
importance matrix
often called:
imatrix
The matrix contains information that helps the quantizer identify which weights are more important for model behavior.
This can improve low-bit quantization choices.
The basic idea is similar to the broader principle:
not every weight deserves the same treatment
This becomes increasingly important at aggressive bit widths.
Requantization
Suppose you have:
Q8 model
and convert it into:
Q4
That is requantization.
This can be worse than starting from the original:
BF16 / FP16
weights and directly producing Q4.
Why?
The Q8 model already contains quantization approximation.
Quantizing that approximation again can introduce additional error.
A safer workflow is generally:
high-precision source
↓
target quantization
rather than:
high precision
↓
Q8
↓
Q4
↓
Q3
unless you understand the quality trade-off.
Model weight quantization vs file format
Another important distinction:
precision
and:
container format
are not the same thing.
A file format tells software how model data and metadata are organized.
Quantization tells you how numerical tensors are represented.
GGUF can contain different tensor types and quantization levels.
So saying:
GGUF model
alone does not tell you how heavily the weights are quantized.
You also need the quantization label.
Model weight quantization vs model architecture
Quantization also does not change:
8B model
into:
4B model
An 8B model quantized to 4 bits still has approximately:
8 billion learned parameters
The parameters are simply represented using fewer bits.
This is a critical distinction.
Quantization reduces:
precision per parameter
not:
number of parameters
Quantization vs pruning
Pruning removes model parameters or structures considered unnecessary.
Conceptually:
Quantization:
same general parameters
→ fewer bits per value
Pruning:
remove some parameters / structures
They are different compression strategies.
They can potentially be combined, but one should not be confused with the other.
Quantization vs distillation
Knowledge distillation trains a smaller model using information from a larger teacher model.
Conceptually:
large teacher
↓
training signal
↓
smaller student
The student actually has a different parameter count or architecture.
Quantization instead normally starts with the same trained model and changes numerical representation.
Again:
quantization
≠
distillation
Quantization vs LoRA
LoRA is a parameter-efficient adaptation technique.
A LoRA adds or trains a small number of low-rank parameters while leaving most base-model weights unchanged.
Quantization reduces numerical precision.
They can be combined.
For example:
quantized base model
+
LoRA adapter
is a common approach to memory-efficient fine-tuning or inference.
What is QLoRA?
QLoRA combines a quantized base model with trainable low-rank adapters.
A well-known QLoRA setup uses:
4-bit base-model weights
+
LoRA adapters
while performing the necessary computation at higher precision.
This dramatically reduces the memory required to fine-tune large models compared with full-parameter training.
The quantized base weights remain frozen while the LoRA parameters are trained.
NF4
You may encounter:
NF4
which means:
NormalFloat 4
NF4 is a 4-bit data type introduced for QLoRA-style quantization and designed around normally distributed neural-network weights.
It is commonly encountered in the bitsandbytes ecosystem.
NF4 should not be confused with:
GGUF Q4_K_M
They are different low-bit representations used in different frameworks and workflows.
LLM.int8()
The bitsandbytes ecosystem also provides:
LLM.int8()
This is not simply naive conversion of every value to ordinary INT8.
The method handles sensitive outlier features differently so that lower-precision matrix multiplication can be used without allowing a small number of large activation values to create excessive error.
This illustrates an important principle:
"8-bit"
often describes a quantization system, not merely a primitive integer type.
Which is better: INT8 or INT4?
There is no universal answer.
INT8 generally uses more memory but provides more numerical resolution.
INT4 uses less memory but requires more aggressive approximation.
A simple conceptual trade-off is:
INT8
→ larger
→ usually safer numerically
INT4
→ smaller
→ potentially more quantization error
But modern 4-bit methods can preserve quality extremely well.
The actual choice depends on:
model
hardware
runtime
memory budget
task
Should you always choose the highest precision that fits?
Not necessarily.
Suppose:
Q8
barely fits in VRAM while:
Q4
leaves several gigabytes free.
That extra memory could be used for:
larger context
KV cache
more concurrent requests
other GPU applications
A slightly lower-precision model may therefore provide a better overall system.
Memory headroom has practical value.
Should you always choose the smallest model file?
Also no.
Aggressive quantization can eventually degrade quality enough to outweigh the memory savings.
A very small quantization may also have worse kernel support or speed on your runtime.
Do not optimize only for:
download size
Optimize for:
quality
+
memory
+
speed
+
compatibility
Choosing quantization for llama.cpp
For llama.cpp and GGUF, start by asking:
How much RAM or VRAM do I have?
Then choose a model and quantization that leave reasonable headroom.
A common practical progression is:
Q4
→ lower memory
Q5
→ somewhat more memory, often higher fidelity
Q6
→ larger again
Q8
→ much closer to high precision in storage cost
The exact effective bits-per-weight and quality characteristics vary between quantization types.
We will examine those in the dedicated GGUF quantization guide.
Choosing quantization for GPU serving
For GPU-oriented serving frameworks, compatibility with optimized kernels matters heavily.
Methods such as:
AWQ
GPTQ
FP8
bitsandbytes
may be supported differently depending on:
GPU architecture
serving framework
model architecture
kernel implementation
Do not choose a quantization format without checking your serving engine.
A model that is theoretically smaller is not useful if your runtime cannot execute it efficiently.
A practical memory example
Suppose you want to run a dense 32B model.
Idealized weight storage:
FP16:
32B × 2 bytes
≈ 64 GB
8-bit:
32B × 1 byte
≈ 32 GB
4-bit:
32B × 0.5 byte
≈ 16 GB
This immediately explains why 4-bit quantization can make a 32B model practical on hardware where FP16 is impossible.
But remember:
16 GB
is an idealized weight-only calculation.
Real inference also needs:
quantization metadata
KV cache
runtime buffers
temporary allocations
For complete sizing, read How Much VRAM Do You Need for Local LLMs?.
Another example: 70B
Idealized FP16 weight memory:
70B × 2 bytes
≈ 140 GB
Idealized 4-bit weight memory:
70B × 0.5 byte
≈ 35 GB
This is why quantization transforms the consumer-hardware feasibility of large models.
But a typical 24 GB GPU still cannot hold an ordinary dense 70B model entirely in VRAM at an idealized 4 bits per parameter.
CPU/GPU hybrid inference or multiple GPUs may still be required.
Why local LLM users care so much about quantization
Cloud AI hides most memory management from the user.
Local AI exposes it directly.
When you run models yourself, you have a fixed amount of:
VRAM
RAM
memory bandwidth
storage
Quantization lets you trade some numerical fidelity for lower resource requirements.
That trade can determine whether a model:
does not run at all
or:
runs comfortably on your computer
This is why quantization is one of the most important concepts in local AI.
A useful mental model
Think of an LLM as a very large collection of numbers.
High precision:
0.1837214
-0.2948172
0.9271831
Quantization replaces those values with a compact representation:
small integer codes
+
scales
+
other metadata
The runtime uses that compact representation to approximate the original model’s computation.
Better quantization methods attempt to decide:
which approximation errors matter
and:
how to minimize them
rather than blindly rounding every number.
What quantization does not change
Quantization does not automatically change:
the tokenizer
training data
model architecture
parameter count
context-window design
knowledge cutoff
instruction tuning
It changes the numerical representation of tensors.
If two files are quantizations of the same original model, they begin from the same underlying learned model but represent its weights differently.
Why two quantizations can produce different answers
LLM decoding is sensitive to probability differences.
Quantization slightly changes internal computations.
That can slightly change token probabilities.
Suppose high precision predicts:
Token A: 40.01%
Token B: 39.99%
A tiny numerical change could reverse them:
Token A: 39.98%
Token B: 40.02%
Once a different token is generated, future context changes.
The entire continuation can diverge.
Therefore two quantizations can produce visibly different responses even when overall model quality remains similar.
Deterministic generation does not eliminate quantization differences
Even with:
temperature = 0
different quantized weights can produce different logits.
The highest-probability token can change.
So deterministic decoding can improve reproducibility within one configuration but does not make different quantization levels mathematically identical.
How should quantization quality be measured?
There is no single perfect metric.
Possible evaluations include:
perplexity
task benchmarks
reasoning tests
coding tests
human evaluation
domain-specific accuracy
output similarity
For local deployment, you should also measure:
tokens per second
VRAM usage
RAM usage
model loading time
context capacity
A slightly smaller quality score might be acceptable if the model becomes dramatically easier to deploy.
Do not compare quantizations using one random prompt
One conversation is not a benchmark.
If you ask:
Write a poem about a cat.
and prefer one answer, that tells you almost nothing about general quantization quality.
A meaningful comparison should use:
repeatable dataset
same model
same runtime
same settings
same hardware
multiple prompts/tasks
This is also why RAMGPT does not label unperformed tests as benchmarks.
Practical beginner recommendation
If you are new to local LLMs, do not start by searching for the mathematically smallest possible quantization.
Instead:
1. Choose a model appropriate for your task
2. Check your available RAM/VRAM
3. Choose a well-supported quantization
4. Leave memory headroom
5. Test real prompts
6. Measure speed
7. Move higher or lower in precision if needed
For llama.cpp users, a mainstream Q4 or Q5 GGUF variant is often easier to experiment with than extreme low-bit variants.
The ideal choice remains model- and hardware-dependent.
Quantization decision tree
A simple decision process:
Does high precision fit comfortably?
|
Yes
↓
Do you need the extra memory?
|
No
↓
Higher precision may be fine
Does high precision NOT fit?
|
Yes
↓
Try 8-bit / 6-bit / 5-bit / 4-bit
↓
Does model now fit comfortably?
|
Yes
↓
Test quality and speed
↓
Choose the best trade-off
Do not treat bit width as a quality ranking.
Treat it as a deployment parameter.
The four concepts to remember
If you remember only four things, remember these.
First:
Quantization reduces numerical precision.
Second:
Lower precision usually reduces model memory.
Third:
4-bit does not mean every runtime value is literally 4 bits.
Fourth:
The best quantization depends on model + runtime + hardware + task.
Those four rules will prevent most common misunderstandings.
Bottom line
LLM quantization makes large models practical by representing model weights using fewer bits.
A simplified progression looks like:
FP32
4 bytes per weight
FP16 / BF16
2 bytes per weight
INT8
about 1 byte per weight
INT4
about 0.5 byte per weight
Real quantized formats require additional metadata and may keep some tensors or computations at higher precision, so actual storage is not always equal to the idealized arithmetic.
Modern methods such as:
AWQ
GPTQ
GGUF quantization
bitsandbytes 4-bit / 8-bit
use more sophisticated strategies than simply rounding every weight.
Quantization can reduce:
disk size
RAM
VRAM
memory bandwidth pressure
and can sometimes improve inference performance.
But aggressive quantization may introduce model-quality loss, and speed depends heavily on whether your runtime and hardware have efficient kernels for that representation.
The right question is therefore not:
What is the smallest quantization?
It is:
What is the lowest-precision representation that gives me the quality, speed, memory use, and compatibility I need?
That is the practical purpose of LLM quantization.