StorySoftware

DeepSeek releases V2, introducing multi-head latent attention to shrink memory use

DeepSeekDeepSeek-V series

DeepSeek released V2, a mixture-of-experts model with 236B total parameters and 21B active per token, using 128K-token context. Multi-head latent attention compresses KV cache memory by 93.3%, reducing training cost 42.5% and raising maximum generation throughput 5.76x versus DeepSeek 67B.

  • Trained on 8.1 trillion tokens via multi-source corpus, followed by supervised fine-tuning and reinforcement learning.
  • Multi-head latent attention achieves efficient inference by compressing key-value cache into latent vector.
  • DeepSeekMoE mechanism enables cost-effective training through sparse computation.
  • 42.5% training cost reduction and 5.76x throughput improvement demonstrate practical computational advantages.
Read the original · Paper

More on DeepSeek-V series

Primer