The brief
- M3 (June 2026): about 428B parameters with 23B active per token and a 1M-token context; weights on Hugging Face.
- M3 scores 80.5% on SWE-Bench Verified, up from M2.5's 80.2%.
- M2 and M2.5 weights are MIT-licensed; M3 uses a MiniMax community licence.
Technical approach
Mixture-of-experts
M3 runs about 23B of its ~428B parameters per token (M2.5: 10B of 230B), cutting cost while keeping quality.
Trained on real-world task environments
The model was extensively trained using reinforcement learning across complex task environments, focusing on coding, tool use, and practical problem solving.
Inference at $1 per hour, 2x faster
M2.5-Lightning: 100 tokens/sec (roughly 40% faster than M2.1); costs 1/10–1/20th of Anthropic, Google, OpenAI.
Downloadable weights
M2 and M2.5 weights are MIT-licensed; M3's use a MiniMax community licence. All are on Hugging Face, with API access via minimax.io.

