StorySoftware

DeepSeek releases V3, a 671B-parameter open-weight mixture-of-experts model

DeepSeekDeepSeek-V series

DeepSeek released V3 as open weights: a mixture-of-experts model with 671B total parameters and 37B active per token, trained on 14.8 trillion tokens. Its technical report put full training at 2.788 million H800 GPU hours, a fraction of what comparable models were thought to cost, with performance close to leading closed models.

  • 671B parameters, 37B active per token; MoE architecture with efficient expert selection
  • Trained on 14.8 trillion diverse, high-quality tokens
  • Training efficiency: 2.788M H800 GPU hours—fraction of comparable models' cost; stable training with no rollbacks
  • Performance comparable to leading closed-source models; outperforms other open-source models
Read the original · Paper

More on DeepSeek-V series

Primer