Milestones
- R1-Lite preview20 Nov 2024Complete.
- DeepSeek-R1 released under MIT licence20 Jan 2025Complete.
- R1-0528 update28 May 2025Complete.
- R1 paper published in Nature17 Sep 2025Complete.
Most important updates
- 17 Sep 2025
- 28 May 2025
- 27 Jan 2025
- 20 Jan 2025
- 20 Nov 2024
Physics limits
- Only checkable skills train wellReasoning is learned by rewarding answers a program can grade, such as maths and code. Where no automatic checker exists (strategy, taste, open research) the signal is weak.
- Better answers need ever longer thinkingReasoning accuracy rises roughly with the logarithm of thinking tokens: each further gain needs several times more working-out, time and compute per answer.
- Trained for likely text, not true textThe model learns to produce plausible words. Facts seen only once or twice in training can't be recalled reliably, so some rate of confident errors is built into the method.
How it works

Learning to reason by trial
R1-Zero was trained only with rewards for right answers on maths and code; long step-by-step reasoning and self-checking emerged on their own.
Grading answers in groups
For each question the model writes several answers, and each is scored against the group's average, so no separate critic model is needed.
A little worked-example data
R1 adds a small set of readable worked examples before RL, fixing R1-Zero's language mixing and hard-to-read reasoning.
Teaching small models
R1's reasoning traces were used to fine-tune small Qwen and Llama models, passing much of the skill to models that run on one GPU.
Update log
Wed 17 Sep 2025
- Minor: PaperResearch
Wed 28 May 2025
- Minor: BlogSoftware
Mon 27 Jan 2025
- Major: PressOther
Mon 20 Jan 2025
- Major: BlogSoftware
Wed 20 Nov 2024
- Minor: BlogSoftware
About DeepSeek
DeepSeek
Open-weight models, V4, R1 reasoning
DeepSeek is a Hangzhou AI lab spun out of the trading fund High-Flyer and led by Liang Wenfeng. Its open V3 and R1 models matched top US models at a fraction of the reported cost, causing a sharp drop in AI chip stocks.
- R1 (January 2025) showed reasoning can emerge from reward-based training with few examples. The paper appeared in Nature.
- V4 arrived in 2026: a preview in April, then V4-Pro (1.6T parameters, 49B active) in August.
- Works under US export limits, training on restricted NVIDIA chips and adapting to Chinese chips.
- Founded
- 20233 yrs
- Headquarters
- China
- Status
- Private
- Valuation
- private
- Works in
- AIOpen-weight models
- Coverage
- 2 programs · 13 updateslatest 30d agochecked 25 Sep
- People
- Liang WenfengFounder and CEO

