AI Fundamentals

What Is LLM Quantization? FP16, INT8 and INT4 Explained

Learn how LLM quantization reduces model memory, what FP16, INT8 and INT4 mean, and how GGUF, AWQ, GPTQ, scales, groups, and quality trade-offs work.

Approximately 25 min read

Download a local LLM and you will quickly encounter labels such as:

FP16
BF16
INT8
INT4
Q4_K_M
AWQ
GPTQ

All of them are related to how model numbers are represented or compressed, but they do not all mean the same thing.

The central idea is quantization.

Quantization reduces the numerical precision used to represent some parts of a neural network, usually its weights.

That can dramatically reduce:

model file size
RAM usage
VRAM usage
memory bandwidth requirements

and may also improve inference performance on hardware and runtimes optimized for the quantized format.

But quantization is not free.

Reducing precision can introduce approximation error.

The practical goal is therefore not:

use the fewest bits possible

It is:

use the lowest precision that still gives
acceptable quality and good hardware performance

This guide explains what that means.

What exactly is being quantized?

A large language model contains billions of learned parameters.

These parameters are numerical values called:

weights

A tiny example might contain values such as:

0.13742
-0.83291
0.00417
1.24931

Real LLMs contain billions of numbers organized into large tensors.

During training, these values are normally represented using floating-point formats.

For deployment, storing every weight at high precision can consume enormous amounts of memory.

Quantization replaces or approximates those high-precision numbers using a smaller numerical representation.

Conceptually:

high-precision weights
        ↓
quantization
        ↓
lower-precision representation

The model becomes smaller.

Why does precision consume memory?

Every model parameter occupies some number of bits.

For a simple theoretical dense model:

FP32
= 32 bits
= 4 bytes per parameter

FP16
= 16 bits
= 2 bytes per parameter

INT8
= 8 bits
= 1 byte per parameter

INT4
= 4 bits
= 0.5 byte per parameter

So the theoretical weight memory is approximately:

parameters × bits per parameter ÷ 8

For an 8-billion-parameter model:

FP16:

8 billion × 16 / 8
≈ 16 GB

At an idealized 8 bits:

8 billion × 8 / 8
≈ 8 GB

At an idealized 4 bits:

8 billion × 4 / 8
≈ 4 GB

This simple calculation explains why 4-bit models are so attractive for local AI.

But real quantized files are usually somewhat more complicated than this ideal calculation.

Why isn’t a 4-bit model exactly 0.5 byte per parameter?

Quantization usually requires metadata.

A group of weights may need additional values such as:

scale
zero point
group metadata
higher-precision tensors
alignment overhead

Some tensors may intentionally remain at higher precision.

Therefore:

4-bit quantization

does not necessarily mean:

exactly 4.000 bits for every parameter in the file

A GGUF quantization, for example, may have an effective bits-per-weight value greater than the nominal number in its label.

That is normal.

FP32

FP32 means:

32-bit floating point

It has historically been a common numerical format for neural-network training and general computation.

Each value uses four bytes.

For large models, this becomes expensive.

A 70B-parameter model stored entirely as FP32 weights would theoretically require roughly:

70 billion × 4 bytes
=
280 GB

for weights alone.

That does not include runtime memory.

For inference, using FP32 weights for a large modern LLM is therefore often unnecessary and impractical.

FP16

FP16 means:

16-bit floating point

It stores each value in two bytes.

Compared with FP32:

FP32
4 bytes

FP16
2 bytes

so idealized weight storage is cut approximately in half.

For a 70B dense model:

70 billion × 2 bytes
≈ 140 GB

again before runtime overhead.

FP16 preserves a floating-point representation but has less numerical precision and range than FP32.

BF16

BF16 means:

bfloat16

Like FP16, it occupies:

16 bits

per value.

But BF16 and FP16 allocate those bits differently between exponent and significand.

BF16 retains an exponent range similar to FP32 while providing less mantissa precision.

This makes BF16 useful in many modern training and inference workloads where supported by the hardware.

From a simple model-storage perspective:

FP16
≈ 2 bytes/parameter

BF16
≈ 2 bytes/parameter

Neither one should normally be described as aggressive low-bit LLM quantization.

They are reduced-precision floating-point formats.

Is FP16 quantization?

The terminology can become fuzzy.

Going from FP32 to FP16 reduces numerical precision, so in a broad mathematical sense it is a precision reduction.

But when people discuss:

LLM quantization

they usually mean more aggressive compression such as:

8-bit
4-bit
3-bit
2-bit

or specialized floating-point formats.

So in practical local-LLM discussion, FP16 and BF16 are usually treated as higher-precision baselines rather than what users mean when they say:

a quantized model

INT8

INT8 uses:

8-bit integer values

instead of storing every weight directly as a 16- or 32-bit floating-point number.

A naive approach might seem simple:

0.134 → 3
0.295 → 7
0.881 → 21

But converting arbitrary floating-point weights to integers requires a mapping.

That mapping is where:

scales
zero points
ranges
groups

become important.

Quantization maps continuous values to fewer levels

Suppose a group of weights lies between:

-1.0 and +1.0

A high-precision format can represent many values within that interval.

An 8-bit integer representation contains only a limited number of integer codes.

The quantization system therefore establishes a mapping between:

real-valued weights

and:

quantized integer codes

Conceptually:

real weight
   ↓
scale / mapping
   ↓
integer code

During inference, computation can use the quantized representation directly or convert information into an appropriate compute representation depending on the kernel.

Scale factor

A common quantization model uses a scale.

In a simplified symmetric example:

quantized value ≈ round(real value / scale)

and approximate reconstruction becomes:

real value ≈ quantized value × scale

The scale determines the spacing between representable values.

This is an approximation.

The original floating-point number may not be exactly representable after quantization.

That difference is called:

quantization error

A tiny example

Suppose:

scale = 0.1

and the original weight is:

0.37

Then:

0.37 / 0.1
=
3.7

Round:

4

Reconstruct:

4 × 0.1
=
0.4

The original value was:

0.37

The approximated value becomes:

0.40

The difference:

0.03

is quantization error.

Real LLM quantization algorithms are far more sophisticated than this small example, but the central idea is the same.

Why fewer bits create more error

Imagine representing values with many possible numerical levels.

High precision gives the model many choices.

Lower precision gives fewer.

Conceptually:

High precision:

|.|.|.|.|.|.|.|.|.|.|.|.|.|.|.|.|


Low precision:

|----|----|----|----|

A weight has to be approximated by one of the available levels.

As precision decreases, the distance between possible representations can become larger.

This increases the risk of changing model behavior.

The challenge is deciding how to quantize weights while minimizing meaningful error.

INT4

INT4 uses approximately four bits to represent quantized values.

Four bits allow:

2^4
=
16

possible bit patterns.

That is dramatically fewer representable states than an 8-bit value.

But the storage advantage is substantial.

Theoretical weight memory becomes approximately one-quarter of a 16-bit baseline:

FP16
16 bits

INT4
4 bits

This is why 4-bit quantization is extremely popular for consumer local LLM inference.

4-bit does not automatically mean terrible quality

A neural network contains many weights, and they do not all contribute equally to output error.

Modern quantization methods take advantage of model structure and weight distributions.

Instead of simply rounding every weight independently using one global scale, techniques can use:

group-wise quantization
per-channel scales
importance information
calibration data
mixed tensor types
error compensation

This can preserve much more model quality than naive rounding.

That is why practical 4-bit LLMs can remain highly useful.

What is group-wise quantization?

Rather than use one scale for an entire enormous tensor, weights can be divided into smaller groups.

For example:

weights 1–128
→ scale A

weights 129–256
→ scale B

weights 257–384
→ scale C

Each group gets a scale that better matches the range of weights inside it.

This can reduce quantization error.

The cost is additional metadata.

Smaller groups generally require more scale values.

So there is a trade-off:

smaller groups
→ potentially better representation
→ more metadata

Per-tensor quantization

The simplest mapping might use one scale for an entire tensor:

Tensor
→ one scale

This has low metadata overhead.

But one scale may poorly represent a tensor whose numerical ranges vary substantially between channels or regions.

Per-channel quantization

A finer strategy can assign separate quantization parameters to channels.

Conceptually:

Channel 1 → scale 1
Channel 2 → scale 2
Channel 3 → scale 3

This gives the quantizer greater flexibility.

Again, the price is additional metadata and implementation complexity.

Symmetric quantization

In symmetric quantization, values are mapped around zero using a symmetric range.

Conceptually:

-negative
   ↓
zero
   ↓
positive

A simplified formula may use:

q = round(x / scale)

with zero naturally corresponding to zero.

This can simplify implementation.

Asymmetric quantization

An asymmetric scheme may use both:

scale
+
zero point

This allows the representable integer range to map to a real-valued range that is not symmetric around zero.

Conceptually:

real value
≈
scale × (quantized value - zero point)

Which strategy is better depends on the data distribution, method, and hardware implementation.

Weight-only quantization

An important distinction is:

weight-only quantization

Here, model weights are stored at low precision while activations or intermediate computation may remain at higher precision.

Conceptually:

Weights
→ INT4

Activations
→ FP16/BF16

AWQ is an example of a weight-only quantization method.

This distinction matters because saying:

4-bit model

does not mean every number used during inference is four bits.

What are activations?

Weights are the learned model parameters.

Activations are intermediate values produced when input passes through the network.

Conceptually:

Input
↓
Layer weights
↓
Activations
↓
Next layer
↓
More activations

During inference, activations are created dynamically.

Weight quantization and activation quantization are therefore separate engineering decisions.

Weight-and-activation quantization

Some methods quantize both:

weights

and:

activations

You may see notation such as:

W8A8

meaning approximately:

8-bit weights
8-bit activations

or:

W4A16

meaning:

4-bit weights
16-bit activations

Exact meaning depends on the framework.

These labels provide more information than simply saying:

4-bit

KV cache quantization is another separate thing

An LLM also maintains the KV cache during autoregressive inference.

That cache can use its own numerical type.

For example, you could theoretically have:

model weights:
4-bit quantized

KV cache:
FP16

or a runtime that supports:

model weights:
4-bit

KV cache:
lower precision

These are independent choices.

Read What Is a KV Cache? How It Works and Why It Uses VRAM for a detailed explanation.

Post-training quantization

A very important category is:

Post-Training Quantization
PTQ

The model is first trained at a higher precision.

Then quantization is applied afterward.

Conceptually:

Train model
↓
High-precision model
↓
Quantize
↓
Deployment model

This is attractive because the extremely expensive original training process does not need to be repeated from scratch just to create a quantized deployment version.

AWQ and GPTQ are examples of post-training quantization methods.

Quantization-aware training

Another approach is:

Quantization-Aware Training
QAT

Here, training or fine-tuning simulates or incorporates quantization effects.

The model can adapt to low-precision behavior during optimization.

Conceptually:

training
+
simulated quantization effects
↓
model adapted for low precision

QAT can produce excellent low-precision results, but it requires additional training computation.

For ordinary users downloading local LLMs, post-training quantized models are much more commonly encountered.

What is AWQ?

AWQ stands for:

Activation-aware Weight Quantization

It is a low-bit weight-only quantization method designed for LLM deployment.

A key idea from AWQ is that model weights are not equally important.

The method uses activation statistics to identify important weight channels and reduce harmful quantization error.

Conceptually:

Calibration activations
↓
identify important channels
↓
adjust scaling
↓
quantize weights

AWQ does not mean:

a model file format

It describes a quantization method.

This distinction is important.

For more detail, read GGUF vs AWQ: What’s the Difference and Which Should You Use?.

What is GPTQ?

GPTQ is another post-training weight quantization method for large transformer models.

It uses approximate second-order information to quantize weights while compensating for quantization error.

The practical outcome is that models can be compressed to low bit widths such as:

4-bit
3-bit

while attempting to preserve accuracy.

GPTQ and AWQ therefore solve similar deployment problems using different quantization strategies.

They should not be treated as simple synonyms.

GGUF vs AWQ vs GPTQ

These terms are frequently mixed together.

A useful mental model is:

GGUF
→ model file/container format used heavily in llama.cpp ecosystem

AWQ
→ quantization method

GPTQ
→ quantization method

A comparison like:

GGUF vs AWQ

is therefore somewhat asymmetric.

In practice, users make that comparison because different runtime ecosystems distribute models using these labels.

But conceptually they exist at different layers.

What are GGUF Q4 and Q5 models?

In the llama.cpp ecosystem, GGUF files can use numerous tensor quantization types.

You may encounter filenames such as:

Q4_0
Q4_K_M
Q5_K_M
Q6_K
Q8_0

These are llama.cpp/GGML-family quantization configurations.

The initial number is useful, but it does not completely describe the effective storage or quality.

For example:

Q4_K_M

does not mean:

every tensor uses exactly four bits

The format can use mixtures of tensor types and additional block metadata.

We will cover these variants in a separate guide:

GGUF Quantization Levels Explained:
Q4, Q5, Q6 and Q8

Q4_K_M is commonly used because it often provides a practical balance among:

model size
memory use
quality
speed

in llama.cpp-style deployments.

But it is not a universal best quantization.

Depending on model, hardware, and quality needs, you may prefer:

Q5
Q6
Q8

or something smaller than Q4.

The correct choice depends on the workload.

Quantization changes model size

Suppose the same model is distributed as:

BF16
Q8
Q6
Q5
Q4

You should expect lower-bit variants to generally require less storage.

This can make previously impossible models practical.

For example:

32B model in BF16

may be far too large for a consumer GPU.

A 4-bit quantized version may fit much more comfortably.

For detailed memory calculations, read How Much VRAM Do You Need for Local LLMs?.

Quantization can change whether a model fits entirely in VRAM

This may be more important than the direct numerical speed difference between quantization formats.

Imagine:

GPU:
24 GB VRAM

Model A:

large quantization
→ does not fit fully
→ CPU/GPU hybrid

Model B:

smaller quantization
→ fits fully in VRAM
→ full GPU inference

Model B can have a major practical performance advantage because the entire placement strategy changes.

Quantization therefore affects:

storage
memory
hardware placement
and potentially speed

Does quantization always make inference faster?

No.

Smaller models reduce memory traffic, which can help performance.

But actual speed also depends on:

hardware
runtime
quantized kernels
model architecture
batch size
context
GPU backend
CPU backend
dequantization overhead

If your hardware has highly optimized kernels for one quantization format, it may perform very well.

Another format might save memory without providing the same speed advantage.

Therefore:

fewer bits
≠
automatically more tokens per second

Benchmark your actual runtime.

Quantized storage vs compute precision

A model may be stored using one precision while computation uses another representation internally.

For example:

stored weights:
4-bit

computation:
FP16/BF16 or specialized integer/mixed kernels

This is implementation-dependent.

Do not assume that a model labeled:

INT4

means the GPU performs every operation using ordinary 4-bit integer arithmetic.

The runtime and kernels decide how data is loaded, unpacked, dequantized, and multiplied.

Why hardware compatibility matters

A quantization format is useful only if your inference runtime and hardware support it efficiently.

For example, a GPU-serving stack may have optimized kernels for specific methods.

A CPU-oriented runtime may prefer a different family of quantization formats.

That is why choosing a quantized model should start with:

runtime
+
hardware

not merely the smallest file available.

Quantization is not compression like ZIP

A ZIP file becomes compressed on disk and is decompressed back into the exact original bytes.

Lossless compression means:

original data
→ compression
→ decompression
→ exact original data

Low-bit LLM quantization is generally different.

It usually changes the numerical representation itself.

Conceptually:

0.37291
→
approximate lower-precision value

The original weight may not be recoverable exactly.

That makes ordinary low-bit quantization a form of lossy numerical compression.

Why can the model still work after billions of values change?

Neural networks have a degree of numerical tolerance and redundancy.

The model’s behavior depends on the collective operation of billions of parameters.

A small approximation in an individual weight does not necessarily cause a meaningful output error.

The real challenge is ensuring that quantization errors do not accumulate in important parts of the network.

Modern methods therefore pay special attention to:

weight distribution
activation behavior
sensitive channels
grouping
scaling
error compensation

Quantization quality is model-dependent

There is no rule saying:

every 4-bit model loses exactly X% quality

The impact depends on:

architecture
parameter count
quantization method
calibration
tensor types
task
evaluation metric

Some models tolerate aggressive quantization better than others.

Some tasks may reveal degradation more clearly.

Therefore, avoid universal quality percentages unless they come from a specific reproducible benchmark.

Larger models can sometimes tolerate quantization surprisingly well

A large model may contain substantial redundancy and capacity.

This means a larger quantized model can sometimes be more useful than a smaller high-precision model within the same memory budget.

For example, the practical choice may be between:

small model at high precision

and:

larger model at lower precision

Which one performs better depends on the models and task.

Parameter count and bit width must be evaluated together.

Calibration data

Some quantization methods use a calibration dataset.

The quantizer feeds representative inputs through the model and observes information such as:

activation ranges
sensitive channels
output error

This helps determine better quantization parameters.

Conceptually:

sample prompts
↓
model activations
↓
quantization statistics
↓
better scaling / quantization decisions

Calibration does not mean retraining the original model from scratch.

Why calibration data matters

If a quantization algorithm uses activation statistics, its decisions depend partly on what it observes.

A calibration dataset should therefore be reasonably representative of expected model behavior.

Poor calibration data can make quantization choices less effective.

How sensitive a method is to calibration data depends on the method.

What is an importance matrix in llama.cpp?

The llama.cpp quantization tooling can optionally use an:

importance matrix

often called:

imatrix

The matrix contains information that helps the quantizer identify which weights are more important for model behavior.

This can improve low-bit quantization choices.

The basic idea is similar to the broader principle:

not every weight deserves the same treatment

This becomes increasingly important at aggressive bit widths.

Requantization

Suppose you have:

Q8 model

and convert it into:

Q4

That is requantization.

This can be worse than starting from the original:

BF16 / FP16

weights and directly producing Q4.

Why?

The Q8 model already contains quantization approximation.

Quantizing that approximation again can introduce additional error.

A safer workflow is generally:

high-precision source
↓
target quantization

rather than:

high precision
↓
Q8
↓
Q4
↓
Q3

unless you understand the quality trade-off.

Model weight quantization vs file format

Another important distinction:

precision

and:

container format

are not the same thing.

A file format tells software how model data and metadata are organized.

Quantization tells you how numerical tensors are represented.

GGUF can contain different tensor types and quantization levels.

So saying:

GGUF model

alone does not tell you how heavily the weights are quantized.

You also need the quantization label.

Model weight quantization vs model architecture

Quantization also does not change:

8B model

into:

4B model

An 8B model quantized to 4 bits still has approximately:

8 billion learned parameters

The parameters are simply represented using fewer bits.

This is a critical distinction.

Quantization reduces:

precision per parameter

not:

number of parameters

Quantization vs pruning

Pruning removes model parameters or structures considered unnecessary.

Conceptually:

Quantization:
same general parameters
→ fewer bits per value

Pruning:
remove some parameters / structures

They are different compression strategies.

They can potentially be combined, but one should not be confused with the other.

Quantization vs distillation

Knowledge distillation trains a smaller model using information from a larger teacher model.

Conceptually:

large teacher
↓
training signal
↓
smaller student

The student actually has a different parameter count or architecture.

Quantization instead normally starts with the same trained model and changes numerical representation.

Again:

quantization
≠
distillation

Quantization vs LoRA

LoRA is a parameter-efficient adaptation technique.

A LoRA adds or trains a small number of low-rank parameters while leaving most base-model weights unchanged.

Quantization reduces numerical precision.

They can be combined.

For example:

quantized base model
+
LoRA adapter

is a common approach to memory-efficient fine-tuning or inference.

What is QLoRA?

QLoRA combines a quantized base model with trainable low-rank adapters.

A well-known QLoRA setup uses:

4-bit base-model weights
+
LoRA adapters

while performing the necessary computation at higher precision.

This dramatically reduces the memory required to fine-tune large models compared with full-parameter training.

The quantized base weights remain frozen while the LoRA parameters are trained.

NF4

You may encounter:

NF4

which means:

NormalFloat 4

NF4 is a 4-bit data type introduced for QLoRA-style quantization and designed around normally distributed neural-network weights.

It is commonly encountered in the bitsandbytes ecosystem.

NF4 should not be confused with:

GGUF Q4_K_M

They are different low-bit representations used in different frameworks and workflows.

LLM.int8()

The bitsandbytes ecosystem also provides:

LLM.int8()

This is not simply naive conversion of every value to ordinary INT8.

The method handles sensitive outlier features differently so that lower-precision matrix multiplication can be used without allowing a small number of large activation values to create excessive error.

This illustrates an important principle:

"8-bit"

often describes a quantization system, not merely a primitive integer type.

Which is better: INT8 or INT4?

There is no universal answer.

INT8 generally uses more memory but provides more numerical resolution.

INT4 uses less memory but requires more aggressive approximation.

A simple conceptual trade-off is:

INT8
→ larger
→ usually safer numerically

INT4
→ smaller
→ potentially more quantization error

But modern 4-bit methods can preserve quality extremely well.

The actual choice depends on:

model
hardware
runtime
memory budget
task

Should you always choose the highest precision that fits?

Not necessarily.

Suppose:

Q8

barely fits in VRAM while:

Q4

leaves several gigabytes free.

That extra memory could be used for:

larger context
KV cache
more concurrent requests
other GPU applications

A slightly lower-precision model may therefore provide a better overall system.

Memory headroom has practical value.

Should you always choose the smallest model file?

Also no.

Aggressive quantization can eventually degrade quality enough to outweigh the memory savings.

A very small quantization may also have worse kernel support or speed on your runtime.

Do not optimize only for:

download size

Optimize for:

quality
+
memory
+
speed
+
compatibility

Choosing quantization for llama.cpp

For llama.cpp and GGUF, start by asking:

How much RAM or VRAM do I have?

Then choose a model and quantization that leave reasonable headroom.

A common practical progression is:

Q4
→ lower memory

Q5
→ somewhat more memory, often higher fidelity

Q6
→ larger again

Q8
→ much closer to high precision in storage cost

The exact effective bits-per-weight and quality characteristics vary between quantization types.

We will examine those in the dedicated GGUF quantization guide.

Choosing quantization for GPU serving

For GPU-oriented serving frameworks, compatibility with optimized kernels matters heavily.

Methods such as:

AWQ
GPTQ
FP8
bitsandbytes

may be supported differently depending on:

GPU architecture
serving framework
model architecture
kernel implementation

Do not choose a quantization format without checking your serving engine.

A model that is theoretically smaller is not useful if your runtime cannot execute it efficiently.

A practical memory example

Suppose you want to run a dense 32B model.

Idealized weight storage:

FP16:
32B × 2 bytes
≈ 64 GB

8-bit:
32B × 1 byte
≈ 32 GB

4-bit:
32B × 0.5 byte
≈ 16 GB

This immediately explains why 4-bit quantization can make a 32B model practical on hardware where FP16 is impossible.

But remember:

16 GB

is an idealized weight-only calculation.

Real inference also needs:

quantization metadata
KV cache
runtime buffers
temporary allocations

For complete sizing, read How Much VRAM Do You Need for Local LLMs?.

Another example: 70B

Idealized FP16 weight memory:

70B × 2 bytes
≈ 140 GB

Idealized 4-bit weight memory:

70B × 0.5 byte
≈ 35 GB

This is why quantization transforms the consumer-hardware feasibility of large models.

But a typical 24 GB GPU still cannot hold an ordinary dense 70B model entirely in VRAM at an idealized 4 bits per parameter.

CPU/GPU hybrid inference or multiple GPUs may still be required.

Why local LLM users care so much about quantization

Cloud AI hides most memory management from the user.

Local AI exposes it directly.

When you run models yourself, you have a fixed amount of:

VRAM
RAM
memory bandwidth
storage

Quantization lets you trade some numerical fidelity for lower resource requirements.

That trade can determine whether a model:

does not run at all

or:

runs comfortably on your computer

This is why quantization is one of the most important concepts in local AI.

A useful mental model

Think of an LLM as a very large collection of numbers.

High precision:

0.1837214
-0.2948172
0.9271831

Quantization replaces those values with a compact representation:

small integer codes
+
scales
+
other metadata

The runtime uses that compact representation to approximate the original model’s computation.

Better quantization methods attempt to decide:

which approximation errors matter

and:

how to minimize them

rather than blindly rounding every number.

What quantization does not change

Quantization does not automatically change:

the tokenizer
training data
model architecture
parameter count
context-window design
knowledge cutoff
instruction tuning

It changes the numerical representation of tensors.

If two files are quantizations of the same original model, they begin from the same underlying learned model but represent its weights differently.

Why two quantizations can produce different answers

LLM decoding is sensitive to probability differences.

Quantization slightly changes internal computations.

That can slightly change token probabilities.

Suppose high precision predicts:

Token A: 40.01%
Token B: 39.99%

A tiny numerical change could reverse them:

Token A: 39.98%
Token B: 40.02%

Once a different token is generated, future context changes.

The entire continuation can diverge.

Therefore two quantizations can produce visibly different responses even when overall model quality remains similar.

Deterministic generation does not eliminate quantization differences

Even with:

temperature = 0

different quantized weights can produce different logits.

The highest-probability token can change.

So deterministic decoding can improve reproducibility within one configuration but does not make different quantization levels mathematically identical.

How should quantization quality be measured?

There is no single perfect metric.

Possible evaluations include:

perplexity
task benchmarks
reasoning tests
coding tests
human evaluation
domain-specific accuracy
output similarity

For local deployment, you should also measure:

tokens per second
VRAM usage
RAM usage
model loading time
context capacity

A slightly smaller quality score might be acceptable if the model becomes dramatically easier to deploy.

Do not compare quantizations using one random prompt

One conversation is not a benchmark.

If you ask:

Write a poem about a cat.

and prefer one answer, that tells you almost nothing about general quantization quality.

A meaningful comparison should use:

repeatable dataset
same model
same runtime
same settings
same hardware
multiple prompts/tasks

This is also why RAMGPT does not label unperformed tests as benchmarks.

Practical beginner recommendation

If you are new to local LLMs, do not start by searching for the mathematically smallest possible quantization.

Instead:

1. Choose a model appropriate for your task
2. Check your available RAM/VRAM
3. Choose a well-supported quantization
4. Leave memory headroom
5. Test real prompts
6. Measure speed
7. Move higher or lower in precision if needed

For llama.cpp users, a mainstream Q4 or Q5 GGUF variant is often easier to experiment with than extreme low-bit variants.

The ideal choice remains model- and hardware-dependent.

Quantization decision tree

A simple decision process:

Does high precision fit comfortably?
        |
       Yes
        ↓
Do you need the extra memory?
        |
       No
        ↓
Higher precision may be fine


Does high precision NOT fit?
        |
       Yes
        ↓
Try 8-bit / 6-bit / 5-bit / 4-bit
        ↓
Does model now fit comfortably?
        |
       Yes
        ↓
Test quality and speed
        ↓
Choose the best trade-off

Do not treat bit width as a quality ranking.

Treat it as a deployment parameter.

The four concepts to remember

If you remember only four things, remember these.

First:

Quantization reduces numerical precision.

Second:

Lower precision usually reduces model memory.

Third:

4-bit does not mean every runtime value is literally 4 bits.

Fourth:

The best quantization depends on model + runtime + hardware + task.

Those four rules will prevent most common misunderstandings.

Bottom line

LLM quantization makes large models practical by representing model weights using fewer bits.

A simplified progression looks like:

FP32
4 bytes per weight

FP16 / BF16
2 bytes per weight

INT8
about 1 byte per weight

INT4
about 0.5 byte per weight

Real quantized formats require additional metadata and may keep some tensors or computations at higher precision, so actual storage is not always equal to the idealized arithmetic.

Modern methods such as:

AWQ
GPTQ
GGUF quantization
bitsandbytes 4-bit / 8-bit

use more sophisticated strategies than simply rounding every weight.

Quantization can reduce:

disk size
RAM
VRAM
memory bandwidth pressure

and can sometimes improve inference performance.

But aggressive quantization may introduce model-quality loss, and speed depends heavily on whether your runtime and hardware have efficient kernels for that representation.

The right question is therefore not:

What is the smallest quantization?

It is:

What is the lowest-precision representation that gives me the quality, speed, memory use, and compatibility I need?

That is the practical purpose of LLM quantization.

Sources and further reading

Continue reading