StorySoftware
Mistral releases Mixtral 8x7B, an open-weight mixture-of-experts model
Mistral released Mixtral 8x7B under Apache 2.0, a mixture-of-experts model with 46.7B total parameters of which only 12.9B are used per token. Mistral said it beat Llama 2 70B on most benchmarks with six times faster inference and matched or beat GPT-3.5, with a 32,000-token context.
- 46.7B total parameters; 12.9B active per token, matching inference speed and cost of 12.9B model.
- 6x faster inference than Llama 2 70B while delivering superior capabilities.
- Instruction-tuned variant achieves 8.30 MT-Bench score, strongest open-source instruction-following.
- 32K token context; proficient in five languages and strong code generation abilities.