StorySoftware

DeepSeek-V3.2-Exp introduces DeepSeek Sparse Attention and cuts API prices

DeepSeekDeepSeek-V series

DeepSeek released the experimental V3.2-Exp with DeepSeek Sparse Attention, which lets the model attend to only the most relevant parts of long inputs, making long contexts cheaper. It performed on par with V3.1-Terminus, and API prices were cut by more than 50%.

  • Sparse Attention achieves fine-grained attention with minimal quality impact, enhancing long-context performance and reducing compute expense.
  • Comparable benchmarks to V3.1-Terminus despite efficiency improvements; V3.1-Terminus temporarily available through API until 15 October.
  • Model weights and technical report (including GPU kernels in TileLang and CUDA) released on Hugging Face.
Read the original · Blog

More on DeepSeek-V series

Primer