StorySoftware

Mistral AI releases Mistral 7B under Apache 2.0

Mistral AIMistral models

Mistral AI released Mistral 7B, a 7.3-billion-parameter model under the permissive Apache 2.0 licence, which it said beat the nearly twice-as-large Llama 2 13B on all benchmarks. It used grouped-query and sliding-window attention to run faster and handle longer text cheaply.

  • 7.3B parameters; outperforms Llama 2 13B (nearly 2x larger) on all benchmarks.
  • Grouped-query attention enables faster inference; sliding-window attention handles longer sequences cheaply.
  • Requires only half the cache memory for 8K sequences versus traditional attention.
  • Fine-tuned chat variant outperforms all 7B models on MT-Bench; widely available across platforms.
Read the original · Blog

More on Mistral models

Primer