StorySoftware
MiniMax releases M3, a 428B-parameter multimodal model with a 1M-token context and downloadable weights
MiniMax released M3, its new flagship and the successor to M2.5. It is a mixture-of-experts model with about 428B parameters, of which about 23B work on each token, and it reads text, images and video natively. Its weights are on Hugging Face under a MiniMax community licence.
- About 428B total parameters, about 23B active per token.
- 1M-token context, using MiniMax Sparse Attention: 9x faster reading and 15x faster writing than M2 at 1M tokens.
- 80.5% on SWE-Bench Verified and 78.1% on MMMU Pro, per MiniMax.
- Weights under a MiniMax community licence, not MIT like M2.5.