StorySoftware
DeepSeek releases V3, a 671B-parameter open-weight mixture-of-experts model
DeepSeek released V3 as open weights: a mixture-of-experts model with 671B total parameters and 37B active per token, trained on 14.8 trillion tokens. Its technical report put full training at 2.788 million H800 GPU hours, a fraction of what comparable models were thought to cost, with performance close to leading closed models.
- 671B parameters, 37B active per token; MoE architecture with efficient expert selection
- Trained on 14.8 trillion diverse, high-quality tokens
- Training efficiency: 2.788M H800 GPU hours—fraction of comparable models' cost; stable training with no rollbacks
- Performance comparable to leading closed-source models; outperforms other open-source models