StorySoftware
DeepSeek releases V2, introducing multi-head latent attention to shrink memory use
DeepSeek released V2, a mixture-of-experts model with 236B total parameters and 21B active per token, using 128K-token context. Multi-head latent attention compresses KV cache memory by 93.3%, reducing training cost 42.5% and raising maximum generation throughput 5.76x versus DeepSeek 67B.
- Trained on 8.1 trillion tokens via multi-source corpus, followed by supervised fine-tuning and reinforcement learning.
- Multi-head latent attention achieves efficient inference by compressing key-value cache into latent vector.
- DeepSeekMoE mechanism enables cost-effective training through sparse computation.
- 42.5% training cost reduction and 5.76x throughput improvement demonstrate practical computational advantages.