Milestones
- DeepSeek-V2 with memory-saving attention (MLA)May 2024Complete.
- DeepSeek-V3 (671B MoE), about 2.8M H800 GPU-hours26 Dec 2024Complete.
- V3.2-Exp introduces DeepSeek Sparse Attention29 Sep 2025Complete.
- DeepSeek-V4 preview open-sourced24 Apr 2026Complete.
- DeepSeek-V4-Pro generally available13 Aug 2026Complete.
Most important updates
- 10 Sep 2026
- 13 Aug 2026
- 24 Apr 2026
- 1 Dec 2025
- 26 Dec 2024
Current obstacles
- Chip export limitsUS rules block DeepSeek from top NVIDIA chips, pushing it to efficiency tricks and less mature Chinese chips.
Physics limits
- Every word re-reads the modelTo write each token the chip must stream all active weights from memory. Memory bandwidth (about 8 TB/s on top GPUs), not arithmetic, caps speed and cost per token.
- Attention cost grows with length squaredIn standard attention every token is compared with every other, so doubling the input quadruples that work. Very long contexts need shortcuts that can miss details.
- All experts must sit in memoryMixture-of-experts saves arithmetic, not memory: a 1-trillion-parameter model needs about 1 TB at 8-bit precision even if only 3% runs per word, so it needs many linked chips.
How it works

Huge but sparse
V3 has 671B parameters but uses about 37B per word, picking from 256 small experts per layer, which keeps serving cheap.
Compressed working memory
Multi-head latent attention squeezes the model's running notes on the input (the KV cache) into a small code, cutting memory for long inputs.
Training in low precision
V3 was trained largely with 8-bit numbers in about 2.8M H800 GPU-hours, a fraction of what comparable US models are thought to use.
Reading selectively
Since V3.2, a quick indexer picks the most relevant earlier tokens for each new one, so cost no longer grows with the square of length.
Update log
Thu 10 Sep
- Minor: BlogSoftware
Thu 13 Aug
- Minor: BlogSoftware
Fri 24 Apr
- Major: BlogSoftware
Mon 1 Dec 2025
- Minor: BlogSoftware
Mon 29 Sep 2025
- Minor: BlogSoftware
Thu 21 Aug 2025
- Minor: BlogSoftware
Thu 26 Dec 2024
- Major: PaperSoftware
Tue 7 May 2024
- Minor: PaperSoftware
About DeepSeek
DeepSeek
Open-weight models, V4, R1 reasoning
DeepSeek is a Hangzhou AI lab spun out of the trading fund High-Flyer and led by Liang Wenfeng. Its open V3 and R1 models matched top US models at a fraction of the reported cost, causing a sharp drop in AI chip stocks.
- R1 (January 2025) showed reasoning can emerge from reward-based training with few examples. The paper appeared in Nature.
- V4 arrived in 2026: a preview in April, then V4-Pro (1.6T parameters, 49B active) in August.
- Works under US export limits, training on restricted NVIDIA chips and adapting to Chinese chips.
- Founded
- 20233 yrs
- Headquarters
- China
- Status
- Private
- Valuation
- private
- Works in
- AIOpen-weight models
- Coverage
- 2 programs · 13 updateslatest 10 Sep 2026checked 25 Sep
- People
- Liang WenfengFounder and CEO

