StorySoftware

Moonshot releases Kimi k1.5, a reasoning model trained with reinforcement learning

Moonshot AIKimi models

Moonshot released Kimi k1.5, a multimodal model trained with reinforcement learning on long chains of reasoning, without tree search or learned value models. Its report claimed o1-level scores: 77.5 on AIME and 96.2 on MATH 500, with a 'long2short' method to improve shorter answers too.

  • RL-trained on long reasoning chains; avoids complex techniques like tree search or value functions
  • Multimodal model with long-context scaling and enhanced policy optimization
  • Benchmark scores: 77.5 AIME, 96.2 MATH 500, 94th percentile Codeforces, 74.9 MathVista
  • Long2short method transfers long-chain reasoning to improve short-chain outputs
Read the original · Paper

More on Kimi models

Primer