StorySoftware
Moonshot releases Kimi K2 Thinking reasoning model
Moonshot released Kimi K2 Thinking, a 1-trillion-parameter mixture-of-experts model that makes 200–300 consecutive tool calls while staying on task. It scored 44.9% on Humanity's Last Exam with tools, has a 256K context, and runs efficiently in 4-bit precision.
- Only 32B parameters activate per token; native INT4 quantization achieves 2× speedup in low-latency mode via Quantization-Aware Training.
- Achieves 99.1% on AIME25 with Python, 71.3% on SWE-bench Verified, and 60.2% on BrowseComp.
- Released under Modified MIT License; optimized for vLLM, SGLang, and KTransformers inference engines.