Mixture of Experts: How MoE Models Activate Only Part of the Network
AI Foundations #19 explains Mixture of Experts models: routers, top-k expert selection, active versus total parameters, load balancing, and why MoE changes inference economics.
Approximately 14 min read · AI Foundations / Lesson 19
Yesterday I learned that quantization asks a memory question:
Can we represent the model’s numbers with fewer bits?
Today the question changes.
Suppose a model has a huge number of parameters. Does every parameter need to participate in every token?
For a normal dense neural network layer, the answer is basically yes: the layer’s weights are there, and each input flows through that layer.
A Mixture of Experts, usually shortened to MoE, changes this idea.
Instead of one large feed-forward path doing the work for every token, an MoE layer contains several alternative feed-forward networks called experts. A small routing mechanism chooses only a few experts for each token.
The mental model I use now is:
many experts exist
-> router scores them for this token
-> choose a small subset
-> run only those experts
-> combine their outputs
That is why an MoE model can have a very large total parameter count while using a much smaller number of active parameters for one token.
The distinction sounds like marketing language until we trace one token through the layer.
Start with the dense transformer we already know
Earlier in AI Foundations, our simplified transformer block looked roughly like this:
input hidden states
-> attention
-> feed-forward network / MLP
-> next transformer block
Real blocks also contain normalization, residual connections, and many architectural details, but this is enough for today’s idea.
In a conventional dense transformer, the feed-forward network is one large learned function. Every token that reaches that layer uses that same set of feed-forward weights.
Conceptually:
token A -> same MLP
token B -> same MLP
token C -> same MLP
The activations are different because the tokens have different hidden states, but the layer’s weight matrices are the same.
Now imagine replacing that one MLP with eight separate MLPs.
expert 1
expert 2
expert 3
expert 4
expert 5
expert 6
expert 7
expert 8
If we ran all eight for every token and averaged them, we would have made the model much more expensive.
MoE gets its advantage from sparse activation: only a small number of experts are selected for each token.
The router decides where a token goes
The router is a learned part of the network.
It takes the token’s current hidden representation and produces a score for each expert.
For eight experts, a toy routing score might look like:
expert 1: 0.3
expert 2: 2.1
expert 3: -0.4
expert 4: 1.7
expert 5: 0.2
expert 6: -0.9
expert 7: 0.8
expert 8: 0.1
If the layer uses top-2 routing, the router chooses the two highest-scoring experts:
expert 2
expert 4
Only those experts process this token.
Another token may have completely different scores and route to:
expert 1
expert 7
So routing is token-dependent.
This is the first big idea:
same MoE layer
+ different token representations
-> potentially different expert paths
A finance analogy that helped me
Imagine an investment company with eight specialist desks:
bonds
energy
technology
foreign exchange
credit risk
derivatives
commodities
real estate
A dense process would send every question to all eight desks and require every desk to produce a full analysis.
An MoE-style process first uses a dispatcher.
A bond-duration question might go mainly to:
bonds + credit risk
A currency-hedging question might go mainly to:
foreign exchange + derivatives
The organization can contain lots of specialist capacity without requiring every specialist to work on every question.
This analogy is useful for the compute pattern, but I have to be careful with it.
Neural-network experts are not guaranteed to become clean human-readable specialists like “the bond expert” or “the French expert.” Researchers can sometimes observe specialization patterns, but the model is free to organize its learned computation in ways that are much less interpretable.
So the safe mental model is:
experts = alternative learned subnetworks
not:
experts = neatly labeled human departments
What does an expert actually contain?
In many transformer MoE architectures, the experts replace or expand the feed-forward / MLP part of selected transformer blocks.
That means an expert is often something conceptually like:
hidden state
-> linear projection
-> activation / gating
-> linear projection
-> output hidden state
The attention mechanism can remain shared and dense while the feed-forward computation becomes sparse.
So a simplified MoE transformer block may look like:
input
-> shared attention
-> router
-> selected expert MLPs
-> weighted combination
-> output
Different models make different architectural choices, so “MoE” does not describe one exact implementation. But sparse expert routing is the common idea.
Top-1 and top-2 routing
Suppose a model has eight experts.
With top-1 routing:
one token -> one selected expert
With top-2 routing:
one token -> two selected experts
The Switch Transformer work is famous for simplifying sparse routing around selecting one expert per token. Mixtral uses a sparse mixture where each token is routed to two experts.
Why use more than one expert?
One intuition is that a token may benefit from combining more than one learned transformation.
But increasing the number of active experts also increases computation.
So there is a tradeoff:
more selected experts
-> potentially more representational capacity per token
-> more expert computation and communication
The best choice is an architectural and training decision, not a universal rule.
The expert outputs are weighted
The router does more than pick IDs.
Its scores can also determine how strongly the selected expert outputs contribute.
For a toy top-2 example, imagine the normalized routing weights are:
expert 2: 0.70
expert 4: 0.30
The two experts produce vectors:
E2(x)
E4(x)
A simplified combination is:
output = 0.70 * E2(x) + 0.30 * E4(x)
The real implementation may differ in details, but this shows why it is called a mixture.
The router chooses a sparse set of functions and combines their contributions.
Total parameters and active parameters are different numbers
This is the part that finally made modern MoE model names less confusing to me.
Imagine a toy model with:
shared parameters: 2 billion
8 experts: 1 billion parameters each
Total expert parameters are:
8 x 1B = 8B
So total model parameters are roughly:
2B shared + 8B experts = 10B total
But if each token activates only two experts, the token does not run through all 8B expert parameters.
Its active path is roughly:
2B shared
+ 2 x 1B selected experts
= 4B active parameters
This is only a toy accounting example. Real architectures share parameters across layers differently, expert sizes vary, and “active parameter” figures reported by model creators can follow architecture-specific definitions.
But the distinction is real:
total parameters = everything stored in the model
active parameters = the subset participating in one token's forward path
That is why a label such as “large total, smaller active” can make sense rather than being mathematically contradictory.
MoE does not make the inactive weights disappear from memory
This point is critical for local inference.
If a model has 100 billion total parameters but only 10 billion are active for one token, it does not mean I only need memory for 10 billion parameters.
The runtime still needs access to whichever experts the router may select.
If all experts are resident on one machine, their weights still consume memory even while most of them are idle for a particular token.
So MoE can reduce compute per token relative to a similarly sized dense model, but the storage problem is closer to the total parameter count.
Very roughly:
memory pressure -> cares a lot about total weights
compute per token -> cares a lot about active weights
This is why quantization and MoE often appear together in local-AI discussions.
Quantization helps shrink all those stored expert weights.
MoE reduces how many expert computations are active for each token.
They solve different parts of the resource problem.
Why MoE can scale model capacity efficiently
Suppose I want to give a model more learned capacity.
With a dense network, adding a much larger feed-forward layer generally increases both:
parameter count
and
compute used for every token
With MoE, I can add more experts while keeping the number selected per token small.
Conceptually:
8 experts, top-2
-> add more stored expert capacity
-> still run only 2 experts per token
This is one reason sparse MoE architectures are attractive for scaling.
They let total parameter capacity grow faster than per-token expert compute.
That does not mean the extra capacity is free. Training, memory, routing, communication, and load balancing all become harder.
But it changes the scaling equation.
The router itself has to learn
How does the model know which experts should receive a token?
The router has trainable parameters too.
During training, gradients update the router along with the rest of the model.
Very roughly:
token representation
-> router scores
-> selected experts
-> model output
-> loss
-> backpropagation
-> update model + routing behavior
So the routing policy is learned from the training objective.
This connects directly back to our earlier lessons on loss, gradient descent, and backpropagation.
The model is not given a hand-written rule such as:
if token is financial, use expert 3
The routing structure emerges from optimization, subject to the architecture and any extra routing or balancing objectives used during training.
A serious problem: what if everyone chooses the same expert?
Imagine eight experts, but the router learns to send almost every token to expert 2.
Then we have created an expensive system where:
expert 2 -> overloaded
experts 1,3,4,5,6,7,8 -> underused
That defeats much of the point.
This is called a load-balancing problem.
Training systems often use additional objectives or routing mechanisms to encourage traffic to be distributed more usefully across experts.
The exact techniques vary by architecture. Some approaches use auxiliary balancing losses; newer designs may use different routing biases or expert structures.
The general engineering requirement is simple:
routing quality is not only about choosing a useful expert
it is also about keeping the expert system usable at scale
Expert capacity can create dropped or rerouted tokens
Hardware also has limits.
During a batch, one expert might suddenly receive many more tokens than another.
A serving or training system has to allocate buffers and schedule that work somehow.
Some MoE designs use an expert capacity concept: an expert has a practical limit on how many tokens it processes in a routing step. Depending on the architecture and implementation, overflow may be dropped, rerouted, padded, or handled with another policy.
This is one reason MoE implementation details matter so much.
Two models can both say “top-2 MoE” while using different strategies for balancing, capacity, shared experts, and communication.
Shared experts are another design
Not every parameter has to sit behind sparse routing.
Some MoE architectures include shared experts that process every token in addition to routed experts.
A simplified picture is:
shared expert -> always active
routed experts -> selected per token
One intuition is that common transformations can live in shared capacity while routed experts handle more specialized computation.
DeepSeekMoE is one example of research exploring finer-grained routed experts together with shared experts.
Again, the important concept is not the branding. It is that MoE architectures can mix:
always-on shared computation
+
sparse token-dependent computation
Why distributed MoE can become a networking problem
So far I imagined all experts on one GPU.
Large MoE models often distribute experts across multiple GPUs or machines.
Now routing a token can require moving its hidden state to the device that owns the selected expert.
Imagine:
GPU 0 holds experts 1-2
GPU 1 holds experts 3-4
GPU 2 holds experts 5-6
GPU 3 holds experts 7-8
A batch of tokens arrives on all GPUs. The router sends different tokens to different experts.
The system may need a communication phase that looks conceptually like:
route tokens to expert-owning GPUs
-> run expert computation
-> send expert outputs back
This often involves all-to-all communication patterns.
That means MoE performance is not only about matrix multiplication speed.
It can depend heavily on:
GPU interconnect bandwidth
network topology
batch size
routing balance
expert placement
communication overlap
A model with lower theoretical active FLOPs can still be difficult to serve efficiently if dispatching tokens between experts becomes the bottleneck.
Why batch size changes the picture
Consider one token at a time.
If that token selects two experts, only those two expert MLPs have useful work.
Now consider a large batch containing many tokens. Across the batch, different tokens may select many different experts.
So even though each token is sparse, the batch as a whole may activate most experts.
This is another reason “only X parameters are active” should not be interpreted as “the rest of the model is irrelevant during serving.”
Per-token sparsity and system-wide utilization are different measurements.
At high concurrency, a runtime may keep many experts busy simultaneously.
MoE is not an ensemble of separate language models
The word “experts” originally made me picture eight complete LLMs voting on the answer.
That is the wrong mental model for transformer MoE.
Usually the experts are components inside one model. They share the rest of the transformer stack, the hidden states, the training objective, and the final output path.
So this:
8 independent LLMs
-> vote
is not what we mean.
A better picture is:
one transformer
-> some layers contain several alternative MLP branches
-> router selects branches token by token
The model remains one trained system.
Does an expert understand a whole topic?
Not necessarily.
A router makes decisions from hidden representations at a particular layer. One token can route differently at different MoE layers as it moves through the network.
So a single sentence might use many different expert combinations across depth.
The specialization can also be statistical rather than semantic. An expert may become useful for patterns that do not have a clean natural-language label.
I now avoid saying things like:
"expert 5 is the math expert"
unless a specific study has actually measured that behavior.
It is more accurate to say:
expert 5 is one learned computation path that the router selects for some hidden states
A token’s journey through an MoE block
Putting the pieces together, one token may travel through an MoE block like this:
1. token already has a hidden-state vector
2. shared attention processes the sequence
3. router scores all available experts
4. choose top-k experts
5. dispatch the token representation to those experts
6. each selected expert runs its MLP
7. weight and combine the expert outputs
8. continue through residual/norm structure
9. pass the result to the next transformer block
The next token may take a different expert route.
And the same token can take different routes at later MoE layers.
That is sparse conditional computation.
What MoE changes for local inference
For someone running models locally, I would summarize the practical effect in four lines:
total parameters -> drives model storage and weight residency
active parameters -> strongly influence per-token expert compute
routing pattern -> determines which expert weights are needed now
memory placement -> determines whether those weights are cheap to access
If all experts fit in fast VRAM, sparse compute can be attractive.
If many expert weights live in system RAM or on slower devices, routing may trigger expensive transfers or memory reads.
That is why MoE offloading is such an active optimization area in local runtimes.
The router may save arithmetic, but it cannot make memory latency disappear.
Quantization and MoE fit together naturally
Yesterday’s lesson now connects to today’s.
Quantization says:
store each weight with fewer bits
MoE says:
use only some expert weights for each token's computation
Together they can make a model with very large total capacity more practical.
But they also create a more complicated runtime:
compressed expert weights
+ routing decisions
+ expert dispatch
+ possible offload
+ possible multi-GPU communication
So an MoE model is not automatically faster than every dense model just because its active parameter count is lower.
Real speed depends on the entire serving path.
What I understand now
My old mental model was:
100B model
-> every token uses all 100B parameters
That works reasonably well as an intuition for a dense model, but it breaks for sparse MoE.
A better model is:
large pool of parameters exists
-> router chooses a small expert subset for each token
-> selected experts do the work
-> outputs are combined
That explains why we need two numbers:
total parameters
active parameters
It also explains the central tradeoff:
more total capacity without proportional per-token expert compute
in exchange for:
routing complexity
load balancing
memory footprint
expert dispatch
communication overhead
MoE is therefore not a compression trick like quantization. It is an architectural choice about conditional computation.
The next step in this course changes the type of data entirely.
So far our language model has mostly consumed tokens derived from text. Modern models can also work with images, audio, and video.
How can those different inputs enter the same transformer-style system?
That leads to multimodal models.
Next: AI Foundations #20 — Multimodal AI: How Text, Images, and Audio Enter One Model.