Primer
Mixture of experts
A neural-network design with many parallel sub-networks ('experts'), where a small router activates only a few of them for each token.
- A router picks, say, 8 of 256 experts per token, so only a small fraction of the model's parameters does work at each step.
- This separates capacity from cost: DeepSeek-V3 has 671 billion parameters but uses only about 37 billion per token.
- Most leading open-weight models, and reportedly several closed ones, are built this way.
- All experts must still sit in memory, usually spread across many chips, and shuttling tokens between them makes communication a bottleneck.