Keep track of advancements in every frontier
DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—DeepSeek-V seriesDeepSeek—

DeepSeek-V series

by DeepSeekCN

ScalingStage 5 of 5

DeepSeek-V4-Pro became generally available on 13 August 2026, reading up to 1M tokens at once.

Updated 10 Sep 2026Checked 25 Sep0 updates this week

Milestones

No announced next step
  1. DeepSeek-V2 with memory-saving attention (MLA)May 2024Complete.
  2. DeepSeek-V3 (671B MoE), about 2.8M H800 GPU-hours26 Dec 2024Complete.
  3. V3.2-Exp introduces DeepSeek Sparse Attention29 Sep 2025Complete.
  4. DeepSeek-V4 preview open-sourced24 Apr 2026Complete.
  5. DeepSeek-V4-Pro generally available13 Aug 2026Complete.

Most important updates

  • 10 Sep 2026
  • 13 Aug 2026
  • 24 Apr 2026
  • 1 Dec 2025
  • 26 Dec 2024

Current obstacles

  • Chip export limitsUS rules block DeepSeek from top NVIDIA chips, pushing it to efficiency tricks and less mature Chinese chips.

Physics limits

  • Every word re-reads the modelTo write each token the chip must stream all active weights from memory. Memory bandwidth (about 8 TB/s on top GPUs), not arithmetic, caps speed and cost per token.
  • Attention cost grows with length squaredIn standard attention every token is compared with every other, so doubling the input quadruples that work. Very long contexts need shortcuts that can miss details.
  • All experts must sit in memoryMixture-of-experts saves arithmetic, not memory: a 1-trillion-parameter model needs about 1 TB at 8-bit precision even if only 3% runs per word, so it needs many linked chips.

How it works

4 parts
JUWELS in Germany, whose booster is built from NVIDIA A100 GPUs, the chip DeepSeek's earlier Fire-Flyer cluster also used
JUWELS in Germany, whose booster is built from NVIDIA A100 GPUs, the chip DeepSeek's earlier Fire-Flyer cluster also usedPhoto: Forschungszentrum Jülich · CC BY-SA 4.0 (opens commons.wikimedia.org)
Experts

Huge but sparse

V3 has 671B parameters but uses about 37B per word, picking from 256 small experts per layer, which keeps serving cheap.

MLA

Compressed working memory

Multi-head latent attention squeezes the model's running notes on the input (the KV cache) into a small code, cutting memory for long inputs.

FP8

Training in low precision

V3 was trained largely with 8-bit numbers in about 2.8M H800 GPU-hours, a fraction of what comparable US models are thought to use.

Sparse

Reading selectively

Since V3.2, a quick indexer picks the most relevant earlier tokens for each new one, so cost no longer grows with the square of length.

Update log

8 updates

Thu 10 Sep

  • Minor: BlogSoftware

Thu 13 Aug

  • Minor: BlogSoftware

Fri 24 Apr

  • Major: BlogSoftware

Mon 1 Dec 2025

  • Minor: BlogSoftware

Mon 29 Sep 2025

  • Minor: BlogSoftware

Thu 21 Aug 2025

  • Minor: BlogSoftware

Thu 26 Dec 2024

  • Major: PaperSoftware

Tue 7 May 2024

  • Minor: PaperSoftware

About DeepSeek

The team behind DeepSeek-V series

DeepSeek

Open-weight models, V4, R1 reasoning

DeepSeek is a Hangzhou AI lab spun out of the trading fund High-Flyer and led by Liang Wenfeng. Its open V3 and R1 models matched top US models at a fraction of the reported cost, causing a sharp drop in AI chip stocks.

  • R1 (January 2025) showed reasoning can emerge from reward-based training with few examples. The paper appeared in Nature.
  • V4 arrived in 2026: a preview in April, then V4-Pro (1.6T parameters, 49B active) in August.
  • Works under US export limits, training on restricted NVIDIA chips and adapting to Chinese chips.
Founded
20233 yrs
Headquarters
China
Status
Private
Valuation
private
Works in
AIOpen-weight models
Coverage
2 programs · 13 updateslatest 10 Sep 2026checked 25 Sep
People
Liang WenfengFounder and CEO
Primer