The brief
- R1 (January 2025) showed reasoning can emerge from reward-based training with few examples. The paper appeared in Nature.
- V4 arrived in 2026: a preview in April, then V4-Pro (1.6T parameters, 49B active) in August.
- Works under US export limits, training on restricted NVIDIA chips and adapting to Chinese chips.
Programs
DeepSeek, two programs
Updated 30d agoDeepSeek-V seriesDeepSeek's main models. V3 (671B parameters, 37B active) used low-precision maths to train cheaply. V3.2 added a faster attention method. V4 grows to 1.6T parameters.ScalingUpdated 30d agoView program →DeepSeek-R1An open-weight reasoning model trained mostly by rewarding correct maths and code answers. Its R1-Zero version learned to think step by step and check itself without worked examples.DeployUpdated 17 Sep 2025View program →
Research · Proto · Pilot · Deploy · Scaling
Peers
Company & funding
- HQ
- CN
- Founded
- 2023
People
- Liang WenfengFounder and CEO

