StorySoftware

DeepSeek releases V4.1-Flash, the first and smallest model of its new V4.1 architecture

DeepSeekDeepSeek-V series

DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with a new causal encoder-decoder design: about 8B parameters are active when reading input and 16B when writing output. It adds image understanding, and DeepSeek says it beats V4-Pro on several benchmarks. From 14 September, V4-Pro API traffic is routed to it at V4.1-Flash's lower prices.

  • Mixture-of-experts means only a small slice of the model runs per token, which keeps it cheap despite its size
  • Its working memory for context (KV cache) takes a quarter of the fast GPU memory of the previous generation, and an eighth of the SSD storage
  • V4-Flash and V4-Flash-Vision-Exp are retired; old model names keep working via routing
  • Weights are on Hugging Face with a technical report; off-peak prices remain half of peak
Read the original · Blog

More on DeepSeek-V series

Primer