StorySoftware

MiniMax releases M3, a 428B-parameter multimodal model with a 1M-token context and downloadable weights

MiniMaxMiniMax M3

MiniMax released M3, its new flagship and the successor to M2.5. It is a mixture-of-experts model with about 428B parameters, of which about 23B work on each token, and it reads text, images and video natively. Its weights are on Hugging Face under a MiniMax community licence.

  • About 428B total parameters, about 23B active per token.
  • 1M-token context, using MiniMax Sparse Attention: 9x faster reading and 15x faster writing than M2 at 1M tokens.
  • 80.5% on SWE-Bench Verified and 78.1% on MMMU Pro, per MiniMax.
  • Weights under a MiniMax community licence, not MIT like M2.5.
Read the original · Blog
Primer