AI Fundamentals

Tensor, Data, Pipeline and Expert Parallelism for LLM Inference

AI Foundations #32 explains the four major ways inference workloads are split across GPUs, what each one replicates or shards, and why communication changes the trade-off.

Approximately 7 min read · AI Foundations / Lesson 32

In AI Foundations #31, we reduced latency by letting a cheaper draft path propose work for the target model.

Now consider a different problem:

One GPU is not enough, or one GPU cannot handle enough requests.

“Use more GPUs” sounds simple, but there are several different ways to do it.

The important question is:

What are we splitting?

That leads to four common forms of inference parallelism:

They can use the same number of GPUs and still behave very differently.

Start with one model on one GPU

Imagine a model as a sequence of layers:

tokens
  |
layer 1
  |
layer 2
  |
layer 3
  |
...
  |
logits

If the weights, KV cache and runtime workspaces fit comfortably on one GPU, one GPU is often the simplest configuration.

Distributed inference adds communication, synchronization and operational complexity.

So the first rule is not:

always use every GPU.

It is:

distribute only when memory, throughput or model structure gives you a reason.

Current vLLM guidance makes the same basic recommendation: if the model fits on a single GPU, distributed inference may be unnecessary.

Tensor parallelism: split a layer across GPUs

Tensor parallelism, or TP, divides large tensor operations inside a layer.

A simplified matrix multiplication is:

Y = XW

If W is too large or expensive for one GPU, we can shard it.

For example:

GPU 0: part of W
GPU 1: part of W
GPU 2: part of W
GPU 3: part of W

Each GPU computes a partial result.

Then the ranks exchange or combine information so the layer produces the correct full result.

The key idea is:

all TP GPUs cooperate on the same forward pass.

What TP buys you

TP can let a model fit when one GPU does not have enough memory.

It can also increase compute capacity for one model replica.

What TP costs

The GPUs must communicate repeatedly during inference.

That means interconnect quality matters.

Four GPUs connected by a fast fabric are not equivalent to four GPUs that can communicate only through a slower path.

TP therefore trades local memory pressure for communication.

Data parallelism: replicate the model, split the requests

Data parallelism, or DP, solves a different problem.

Instead of splitting one copy of the model, we keep multiple replicas:

GPU 0: full model <- requests A, B
GPU 1: full model <- requests C, D
GPU 2: full model <- requests E, F
GPU 3: full model <- requests G, H

Each replica processes different requests.

For dense models, the mental model is close to horizontal scaling.

What DP buys you

DP is primarily a throughput strategy.

If one model replica already fits on a GPU, adding replicas can serve more independent requests.

What DP costs

Weights are replicated.

If the model consumes 20 GB per replica, four DP ranks need roughly four copies of those weights across the GPUs.

Each rank also has its own request state and KV cache.

Current vLLM documentation notes another consequence: each DP engine has an independent KV cache, so routing can affect prefix-cache reuse.

A request sent to the “wrong” replica may miss a prefix that exists on another replica.

TP versus DP

The cleanest distinction is:

TP: several GPUs cooperate on one model replica
DP: several model replicas process different request batches

If a model does not fit on one GPU, DP alone does not solve the problem because every replica still needs the full model.

If the model does fit and the goal is more independent request throughput, DP can be attractive.

Pipeline parallelism: split layers into stages

Pipeline parallelism, or PP, splits the model by depth.

Instead of each GPU holding part of every large tensor, different GPUs hold different groups of layers.

For example:

GPU 0: layers 0-19
GPU 1: layers 20-39
GPU 2: layers 40-59
GPU 3: layers 60-79

A request moves from stage to stage.

This is conceptually similar to a factory line.

Why PP helps

Pipeline parallelism can distribute a model across GPUs or nodes when tensor parallelism across the whole machine would be awkward or communication-heavy.

Current vLLM scaling guidance recommends combining TP and PP for some multi-node cases, often using TP within a node and PP across nodes.

The pipeline bubble

A pipeline is not automatically fully busy.

At the beginning, later stages may wait for earlier stages to produce work.

At the end, earlier stages may become idle while later stages finish.

Scheduling multiple microbatches can fill the pipeline more effectively, but pipeline utilization becomes another systems problem.

Expert parallelism: split the MoE experts

Mixture-of-Experts models add a special opportunity.

A dense feed-forward layer applies the same large block to every token.

An MoE layer has many expert blocks, while a router selects only some experts for each token.

Conceptually:

token
  |
router
 / | \
E2 E7 E11
 \ | /
 combined result

Expert parallelism, or EP, distributes those experts across ranks.

For example:

GPU 0: experts 0-7
GPU 1: experts 8-15
GPU 2: experts 16-23
GPU 3: experts 24-31

Now routed tokens must travel to the GPUs that own their selected experts.

That creates all-to-all communication.

Why EP is different from TP

TP shards tensor computation across ranks.

EP follows the model’s sparse structure: different experts live on different ranks.

For an MoE model, that can improve locality and avoid treating every expert layer like a dense layer.

Current vLLM documentation supports expert parallelism and describes EP size as related to the TP and DP layout used by the deployment.

Communication is part of the model-serving design

It is tempting to think only about parameter memory:

model is 80 GB
4 GPUs have 24 GB each
80 < 96
therefore it fits

That arithmetic is necessary but incomplete.

Distributed inference also depends on:

The same four GPUs can produce very different results under different parallel layouts.

Combining parallelism

Real deployments often combine strategies.

Suppose there are eight GPUs.

One possible layout is:

TP = 2
DP = 4
total GPUs = 8

Each model replica uses two GPUs with tensor parallelism.

There are four such replicas processing different request groups.

For an MoE model with expert parallelism enabled, expert layers may be distributed across a larger EP group while attention layers follow the TP/DP structure.

This is why “we run on eight GPUs” says surprisingly little about the actual architecture.

A decision tree

Use this as a first mental model.

The model fits on one GPU and you need more throughput

Consider data parallel replicas.

The model does not fit on one GPU but fits inside one multi-GPU node

Consider tensor parallelism.

The model is too large for one node

Consider tensor parallelism inside nodes plus pipeline parallelism across nodes.

The model is MoE

Evaluate expert parallelism, often together with data parallelism.

These are starting points, not universal rules.

Why more GPUs can be slower

Distributed execution adds coordination.

If the original workload is small, communication can dominate.

For example:

1 GPU:
compute 4 ms

2-way TP:
compute 2.5 ms
communication 2.5 ms
total 5 ms

Those numbers are invented for illustration, but the principle is real.

Parallelism improves performance only when the saved compute or memory pressure is worth the extra communication and scheduling cost.

The most useful measurement question

When comparing two deployments, do not ask only:

How many tokens per second?

Also ask:

What work is each GPU doing, and what must cross the GPU boundary every step?

That question often explains why a supposedly larger system is not faster.

One compact summary

Tensor parallelism
  split tensors inside layers
  same request spans multiple GPUs

Pipeline parallelism
  split layer ranges
  request moves through stages

Data parallelism
  replicate the model
  different replicas handle different requests

Expert parallelism
  split MoE experts
  routed tokens travel to expert owners

The deeper lesson is that GPUs are not just a pool of memory.

A distributed LLM is a communication topology executing a model.

Once you know what is replicated, what is sharded and what must synchronize, multi-GPU inference becomes much easier to reason about.

Sources and further reading

Continue reading