StorySoftware

Mistral releases Mixtral 8x7B, an open-weight mixture-of-experts model

Mistral AIMistral models

Mistral released Mixtral 8x7B under Apache 2.0, a mixture-of-experts model with 46.7B total parameters of which only 12.9B are used per token. Mistral said it beat Llama 2 70B on most benchmarks with six times faster inference and matched or beat GPT-3.5, with a 32,000-token context.

  • 46.7B total parameters; 12.9B active per token, matching inference speed and cost of 12.9B model.
  • 6x faster inference than Llama 2 70B while delivering superior capabilities.
  • Instruction-tuned variant achieves 8.30 MT-Bench score, strongest open-source instruction-following.
  • 32K token context; proficient in five languages and strong code generation abilities.
Read the original · Blog

More on Mistral models

Primer