StorySoftware
DeepSeek releases V4.1-Flash, the first and smallest model of its new V4.1 architecture
DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with a new causal encoder-decoder design: about 8B parameters are active when reading input and 16B when writing output. It adds image understanding, and DeepSeek says it beats V4-Pro on several benchmarks. From 14 September, V4-Pro API traffic is routed to it at V4.1-Flash's lower prices.
- Mixture-of-experts means only a small slice of the model runs per token, which keeps it cheap despite its size
- Its working memory for context (KV cache) takes a quarter of the fast GPU memory of the previous generation, and an eighth of the SSD storage
- V4-Flash and V4-Flash-Vision-Exp are retired; old model names keep working via routing
- Weights are on Hugging Face with a technical report; off-peak prices remain half of peak