StorySoftware

Meta releases Llama 4 Scout and Maverick mixture-of-experts models

MetaLlama

Meta released Llama 4 Scout and Maverick, its first natively multimodal open-weight models using mixture-of-experts architecture. Scout has 17B active parameters (109B total) with 10-million-token context; Maverick has 17B active (400B total). Both distilled from a 288B Behemoth teacher model and beat GPT-4o and Gemini 2.0 on multiple benchmarks.

  • Scout: 109B total, 17B active, 10M token context, fits on single H100 GPU with Int4 quantization
  • Maverick: 400B total, 17B active, runs on H100 DGX; outperforms GPT-4o and Gemini 2.0, matches DeepSeek V3
  • Native multimodality via early fusion; support up to 8 images in post-training
  • Distilled from Llama 4 Behemoth (288B, still training) used as teacher model
Read the original · Blog

More on Llama

Primer