Tensor, Data, Pipeline and Expert Parallelism for LLM Inference
AI Foundations #32 explains the four major ways inference workloads are split across GPUs, what each one replicates or shards, and why communication changes the trade-off.
Approximately 7 min read · AI Foundations / Lesson 32
In AI Foundations #31, we reduced latency by letting a cheaper draft path propose work for the target model.
Now consider a different problem:
One GPU is not enough, or one GPU cannot handle enough requests.
“Use more GPUs” sounds simple, but there are several different ways to do it.
The important question is:
What are we splitting?
That leads to four common forms of inference parallelism:
- tensor parallelism;
- pipeline parallelism;
- data parallelism;
- expert parallelism.
They can use the same number of GPUs and still behave very differently.
Start with one model on one GPU
Imagine a model as a sequence of layers:
tokens
|
layer 1
|
layer 2
|
layer 3
|
...
|
logits
If the weights, KV cache and runtime workspaces fit comfortably on one GPU, one GPU is often the simplest configuration.
Distributed inference adds communication, synchronization and operational complexity.
So the first rule is not:
always use every GPU.
It is:
distribute only when memory, throughput or model structure gives you a reason.
Current vLLM guidance makes the same basic recommendation: if the model fits on a single GPU, distributed inference may be unnecessary.
Tensor parallelism: split a layer across GPUs
Tensor parallelism, or TP, divides large tensor operations inside a layer.
A simplified matrix multiplication is:
Y = XW
If W is too large or expensive for one GPU, we can shard it.
For example:
GPU 0: part of W
GPU 1: part of W
GPU 2: part of W
GPU 3: part of W
Each GPU computes a partial result.
Then the ranks exchange or combine information so the layer produces the correct full result.
The key idea is:
all TP GPUs cooperate on the same forward pass.
What TP buys you
TP can let a model fit when one GPU does not have enough memory.
It can also increase compute capacity for one model replica.
What TP costs
The GPUs must communicate repeatedly during inference.
That means interconnect quality matters.
Four GPUs connected by a fast fabric are not equivalent to four GPUs that can communicate only through a slower path.
TP therefore trades local memory pressure for communication.
Data parallelism: replicate the model, split the requests
Data parallelism, or DP, solves a different problem.
Instead of splitting one copy of the model, we keep multiple replicas:
GPU 0: full model <- requests A, B
GPU 1: full model <- requests C, D
GPU 2: full model <- requests E, F
GPU 3: full model <- requests G, H
Each replica processes different requests.
For dense models, the mental model is close to horizontal scaling.
What DP buys you
DP is primarily a throughput strategy.
If one model replica already fits on a GPU, adding replicas can serve more independent requests.
What DP costs
Weights are replicated.
If the model consumes 20 GB per replica, four DP ranks need roughly four copies of those weights across the GPUs.
Each rank also has its own request state and KV cache.
Current vLLM documentation notes another consequence: each DP engine has an independent KV cache, so routing can affect prefix-cache reuse.
A request sent to the “wrong” replica may miss a prefix that exists on another replica.
TP versus DP
The cleanest distinction is:
TP: several GPUs cooperate on one model replica
DP: several model replicas process different request batches
If a model does not fit on one GPU, DP alone does not solve the problem because every replica still needs the full model.
If the model does fit and the goal is more independent request throughput, DP can be attractive.
Pipeline parallelism: split layers into stages
Pipeline parallelism, or PP, splits the model by depth.
Instead of each GPU holding part of every large tensor, different GPUs hold different groups of layers.
For example:
GPU 0: layers 0-19
GPU 1: layers 20-39
GPU 2: layers 40-59
GPU 3: layers 60-79
A request moves from stage to stage.
This is conceptually similar to a factory line.
Why PP helps
Pipeline parallelism can distribute a model across GPUs or nodes when tensor parallelism across the whole machine would be awkward or communication-heavy.
Current vLLM scaling guidance recommends combining TP and PP for some multi-node cases, often using TP within a node and PP across nodes.
The pipeline bubble
A pipeline is not automatically fully busy.
At the beginning, later stages may wait for earlier stages to produce work.
At the end, earlier stages may become idle while later stages finish.
Scheduling multiple microbatches can fill the pipeline more effectively, but pipeline utilization becomes another systems problem.
Expert parallelism: split the MoE experts
Mixture-of-Experts models add a special opportunity.
A dense feed-forward layer applies the same large block to every token.
An MoE layer has many expert blocks, while a router selects only some experts for each token.
Conceptually:
token
|
router
/ | \
E2 E7 E11
\ | /
combined result
Expert parallelism, or EP, distributes those experts across ranks.
For example:
GPU 0: experts 0-7
GPU 1: experts 8-15
GPU 2: experts 16-23
GPU 3: experts 24-31
Now routed tokens must travel to the GPUs that own their selected experts.
That creates all-to-all communication.
Why EP is different from TP
TP shards tensor computation across ranks.
EP follows the model’s sparse structure: different experts live on different ranks.
For an MoE model, that can improve locality and avoid treating every expert layer like a dense layer.
Current vLLM documentation supports expert parallelism and describes EP size as related to the TP and DP layout used by the deployment.
Communication is part of the model-serving design
It is tempting to think only about parameter memory:
model is 80 GB
4 GPUs have 24 GB each
80 < 96
therefore it fits
That arithmetic is necessary but incomplete.
Distributed inference also depends on:
- which tensors are sharded;
- how often ranks synchronize;
- how many bytes move per token;
- whether communication overlaps computation;
- whether the interconnect is NVLink, PCIe or network fabric;
- whether the workload is prefill-heavy or decode-heavy;
- whether the model is dense or MoE;
- how KV cache is placed;
- how requests are routed.
The same four GPUs can produce very different results under different parallel layouts.
Combining parallelism
Real deployments often combine strategies.
Suppose there are eight GPUs.
One possible layout is:
TP = 2
DP = 4
total GPUs = 8
Each model replica uses two GPUs with tensor parallelism.
There are four such replicas processing different request groups.
For an MoE model with expert parallelism enabled, expert layers may be distributed across a larger EP group while attention layers follow the TP/DP structure.
This is why “we run on eight GPUs” says surprisingly little about the actual architecture.
A decision tree
Use this as a first mental model.
The model fits on one GPU and you need more throughput
Consider data parallel replicas.
The model does not fit on one GPU but fits inside one multi-GPU node
Consider tensor parallelism.
The model is too large for one node
Consider tensor parallelism inside nodes plus pipeline parallelism across nodes.
The model is MoE
Evaluate expert parallelism, often together with data parallelism.
These are starting points, not universal rules.
Why more GPUs can be slower
Distributed execution adds coordination.
If the original workload is small, communication can dominate.
For example:
1 GPU:
compute 4 ms
2-way TP:
compute 2.5 ms
communication 2.5 ms
total 5 ms
Those numbers are invented for illustration, but the principle is real.
Parallelism improves performance only when the saved compute or memory pressure is worth the extra communication and scheduling cost.
The most useful measurement question
When comparing two deployments, do not ask only:
How many tokens per second?
Also ask:
What work is each GPU doing, and what must cross the GPU boundary every step?
That question often explains why a supposedly larger system is not faster.
One compact summary
Tensor parallelism
split tensors inside layers
same request spans multiple GPUs
Pipeline parallelism
split layer ranges
request moves through stages
Data parallelism
replicate the model
different replicas handle different requests
Expert parallelism
split MoE experts
routed tokens travel to expert owners
The deeper lesson is that GPUs are not just a pool of memory.
A distributed LLM is a communication topology executing a model.
Once you know what is replicated, what is sharded and what must synchronize, multi-GPU inference becomes much easier to reason about.