Primer

Mixture of experts

A neural-network design with many parallel sub-networks ('experts'), where a small router activates only a few of them for each token.

  • A router picks, say, 8 of 256 experts per token, so only a small fraction of the model's parameters does work at each step.
  • This separates capacity from cost: DeepSeek-V3 has 671 billion parameters but uses only about 37 billion per token.
  • Most leading open-weight models, and reportedly several closed ones, are built this way.
  • All experts must still sit in memory, usually spread across many chips, and shuttling tokens between them makes communication a bottleneck.
Primer