StorySoftware
DeepSeek-V3.2-Exp introduces DeepSeek Sparse Attention and cuts API prices
DeepSeek released the experimental V3.2-Exp with DeepSeek Sparse Attention, which lets the model attend to only the most relevant parts of long inputs, making long contexts cheaper. It performed on par with V3.1-Terminus, and API prices were cut by more than 50%.
- Sparse Attention achieves fine-grained attention with minimal quality impact, enhancing long-context performance and reducing compute expense.
- Comparable benchmarks to V3.1-Terminus despite efficiency improvements; V3.1-Terminus temporarily available through API until 15 October.
- Model weights and technical report (including GPU kernels in TileLang and CUDA) released on Hugging Face.