Latency and Throughput: Why LLM Speed Has More Than One Number
AI Foundations #27 explains time to first token, time per output token, prefill, decode, batching, throughput, context length, and why one tokens-per-second number can hide the real user experience.
Approximately 7 min read · AI Foundations / Lesson 27
In AI Foundations #26, we moved from model quality to production behavior.
Now we need one of the most misunderstood production concepts:
How fast is an LLM?
A single number such as 120 tokens/s looks precise, but it may describe only one part of the request.
A user experiences at least two different phases:
read the prompt
-> produce the first token
-> continue generating tokens
Those phases stress hardware differently.
That is why LLM performance needs more than one metric.
Prefill comes first
Suppose the prompt contains 8,000 tokens.
Before the model can generate a reply, it must process those input tokens.
That phase is commonly called prefill.
Conceptually:
prompt tokens
-> forward computation over the prompt
-> KV cache is populated
-> generation can begin
Prefill is usually highly parallel.
A GPU can process many prompt tokens together, so prefill often looks more like a large matrix-compute workload.
Decode is different
After prefill, the model begins generating.
Autoregressive generation works roughly like:
generate token 1
-> update state
-> generate token 2
-> update state
-> generate token 3
Each new token depends on the earlier sequence.
This phase is usually called decode.
Decode has less parallel work per sequence than prefill and often becomes limited by memory movement, cache access, or small repeated kernels.
So a GPU can be excellent at prefill and less impressive at decode.
Time to first token
A useful latency metric is TTFT:
TTFT = time from request arrival to first generated token
TTFT can include queueing, tokenization, scheduler delay, prompt processing, prefill, and network overhead.
For an interactive chat system, TTFT strongly affects perceived responsiveness.
A reply that eventually generates quickly can still feel slow if the user waits several seconds before anything appears.
Time per output token
After the first token, users care about how quickly the rest arrive.
A common metric is TPOT:
TPOT = average time between generated output tokens
Its inverse is approximately decode throughput for one active sequence.
For example:
TPOT = 10 ms/token
-> roughly 100 tokens/s
under a simplified single-stream view.
End-to-end latency combines both phases
Suppose:
TTFT = 1.2 s
output = 200 tokens
TPOT = 12 ms/token
A rough end-to-end time is:
1.2 s
+ 199 x 0.012 s
≈ 3.6 s
The exact accounting depends on how the first token is counted.
The durable mental model is:
request latency
≈ wait for first token
+ time to stream the remaining output
Why prompt length changes the experience
A longer prompt gives prefill more work.
So a 1K prompt and a 64K prompt can have very different TTFT even if decode speed afterward is similar.
Long context also expands the KV cache.
That can increase memory pressure and change which kernels or cache layouts are practical.
So “the model runs at 100 tokens/s” is incomplete unless we know the context size.
Throughput is not the same as latency
Latency asks:
How long does one request take?
Throughput asks:
How much total work can the system complete per unit time?
A server may improve total throughput by batching many users together.
For example:
one request:
80 output tokens/s
eight concurrent requests:
possibly lower speed per request
but much higher total tokens/s
The exact values depend on hardware, model, scheduler, and workload.
The important point is that optimizing total server throughput can make an individual request slower.
Batching helps the GPU stay busy
A GPU is built for parallel work.
If one decode stream does not create enough parallel computation, serving multiple sequences together can improve hardware utilization.
This is one reason modern inference engines use continuous batching.
Instead of waiting for a fixed batch to finish, the scheduler can add and remove requests as sequences arrive and complete.
Conceptually:
request A decoding
request B decoding
request C finishes
request D joins
The batch changes over time.
This can keep the GPU busier than a rigid one-request-at-a-time loop.
But batching creates queueing
Batching is not free.
If a request waits for scheduler capacity, its TTFT can increase.
So a production system often balances:
GPU utilization
vs
per-user latency
At low traffic, there may be little reason to delay a request.
At high traffic, batching can be necessary to keep the system efficient.
Prefill can interfere with decode
A giant prompt can create a long prefill operation.
If the runtime schedules it poorly, users who are already decoding may stall while the large prefill monopolizes compute.
This is one reason modern serving research studies techniques such as chunked prefill.
Instead of processing one enormous prompt as one indivisible block, the runtime can interleave pieces of prefill with ongoing decode work.
The goal is not merely higher benchmark throughput.
It is more predictable service for mixed workloads.
Memory changes with concurrency
Each active sequence needs state.
For Transformer inference, KV cache is a major component.
So higher concurrency can create:
more active sequences
-> more KV cache
-> more memory pressure
Eventually the system can hit a capacity wall even when raw compute is still available.
This is why serving systems such as vLLM focus heavily on KV-cache memory management.
The performance problem is partly scheduling and partly memory allocation.
Average tokens/s can hide a bad tail
Suppose most requests are fast but a few wait a long time.
An average may still look good.
Production systems therefore often inspect percentiles:
p50
p95
p99
for metrics such as TTFT and end-to-end latency.
A p99 TTFT spike can indicate that a small group of users is waiting much longer than the median user.
This connects directly back to monitoring in AI Foundations #26.
One benchmark request is not a server benchmark
Running one prompt, one response, and one GPU can be useful for hardware comparison.
But it does not answer:
How many users can this server support?
What happens at 8 concurrent requests?
What happens with mixed 1K and 64K prompts?
What happens when outputs are long?
What happens when KV cache fills?
Those are serving questions.
A proper server test needs a workload distribution, not just one prompt.
A useful performance report separates the metrics
For a local single-user benchmark, report at least:
model
quantization
hardware
runtime
prompt tokens
output tokens
context limit
prefill tokens/s
decode tokens/s
peak memory
For a serving benchmark, add:
concurrency
request rate
TTFT
TPOT
end-to-end latency
total throughput
p50/p95/p99
Without that context, two “tokens/s” numbers may be measuring different things.
Why a slower model can feel faster
Imagine two systems.
System A:
TTFT = 4.0 s
decode = 150 tok/s
System B:
TTFT = 0.5 s
decode = 100 tok/s
For a short answer, System B may feel much faster because it starts responding almost immediately.
For a very long answer, System A may eventually finish first.
So “faster” depends on the workload.
The cumulative Foundations map
We can now extend the sequence:
parameters and weights
-> tensors
-> forward pass
-> training
-> loss
-> gradient descent
-> backpropagation
-> embeddings
-> tokenization
-> attention
-> Transformer
-> pretraining
-> fine-tuning
-> inference
-> sampling
-> context and KV cache
-> quantization
-> mixture of experts
-> multimodal models
-> reasoning and reinforcement learning
-> evaluation
-> retrieval
-> tools and agents
-> alignment and guardrails
-> monitoring and drift
-> latency and throughput
This final step adds the operational cost of making all the previous pieces run for real users.
The main takeaway
LLM speed is not one number.
The durable mental model is:
prefill
-> how quickly the prompt is processed
TTFT
-> how long the user waits for the first token
decode / TPOT
-> how quickly output continues
throughput
-> how much total work the server completes
tail latency
-> how bad the slow requests become
A benchmark becomes useful only when it tells us which of those it measured.
The next time you see “150 tokens/s,” the right follow-up question is:
At what context, for which phase, with how many simultaneous requests?