HomeFull speed ahead. Keep track of advancements in every frontier
DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—

DeepSeek-R1

by DeepSeekCN

DeployStage 4 of 5

Reasoning is now built into the V-series models, since V3.1.

Updated 17 Sep 2025Checked 25 Sep0 updates this week

Milestones

No announced next step
  1. R1-Lite preview20 Nov 2024Complete.
  2. DeepSeek-R1 released under MIT licence20 Jan 2025Complete.
  3. R1-0528 update28 May 2025Complete.
  4. R1 paper published in Nature17 Sep 2025Complete.

Most important updates

  • 17 Sep 2025
  • 28 May 2025
  • 27 Jan 2025
  • 20 Jan 2025
  • 20 Nov 2024

Physics limits

  • Only checkable skills train wellReasoning is learned by rewarding answers a program can grade, such as maths and code. Where no automatic checker exists (strategy, taste, open research) the signal is weak.
  • Better answers need ever longer thinkingReasoning accuracy rises roughly with the logarithm of thinking tokens: each further gain needs several times more working-out, time and compute per answer.
  • Trained for likely text, not true textThe model learns to produce plausible words. Facts seen only once or twice in training can't be recalled reliably, so some rate of confident errors is built into the method.

How it works

4 parts
An aisle of the Sierra GPU supercomputer (US): reasoning models like R1 are trained with reinforcement learning on clusters like this
An aisle of the Sierra GPU supercomputer (US): reasoning models like R1 are trained with reinforcement learning on clusters like thisPhoto: U.S. Department of Energy · Public domain (opens commons.wikimedia.org)
Pure RL

Learning to reason by trial

R1-Zero was trained only with rewards for right answers on maths and code; long step-by-step reasoning and self-checking emerged on their own.

GRPO

Grading answers in groups

For each question the model writes several answers, and each is scored against the group's average, so no separate critic model is needed.

Polish

A little worked-example data

R1 adds a small set of readable worked examples before RL, fixing R1-Zero's language mixing and hard-to-read reasoning.

Distil

Teaching small models

R1's reasoning traces were used to fine-tune small Qwen and Llama models, passing much of the skill to models that run on one GPU.

Update log

5 updates

Wed 17 Sep 2025

  • Minor: PaperResearch

Wed 28 May 2025

  • Minor: BlogSoftware

Mon 27 Jan 2025

  • Major: PressOther

Mon 20 Jan 2025

  • Major: BlogSoftware

Wed 20 Nov 2024

  • Minor: BlogSoftware

About DeepSeek

The team behind DeepSeek-R1

DeepSeek

Open-weight models, V4, R1 reasoning

DeepSeek is a Hangzhou AI lab spun out of the trading fund High-Flyer and led by Liang Wenfeng. Its open V3 and R1 models matched top US models at a fraction of the reported cost, causing a sharp drop in AI chip stocks.

  • R1 (January 2025) showed reasoning can emerge from reward-based training with few examples. The paper appeared in Nature.
  • V4 arrived in 2026: a preview in April, then V4-Pro (1.6T parameters, 49B active) in August.
  • Works under US export limits, training on restricted NVIDIA chips and adapting to Chinese chips.
Founded
20233 yrs
Headquarters
China
Status
Private
Valuation
private
Works in
AIOpen-weight models
Coverage
2 programs · 13 updateslatest 30d agochecked 25 Sep
People
Liang WenfengFounder and CEO
Primer