StorySoftware
Meta releases Llama 3.1 including a 405B-parameter open model
Meta released Llama 3.1 in 8B, 70B and 405B sizes with 128K-token context and eight-language support. The 405B model, trained on over 16,000 H100 GPUs and over 15 trillion tokens, rivals GPT-4o and Claude 3.5 Sonnet; licensing now allows developers to use Llama outputs to train other models.
- 405B model is the first openly available model competitive with leading closed-source models.
- Trained via iterative post-training with supervised fine-tuning and direct preference optimization.
- License change enables synthetic data generation and model distillation workflows previously unavailable in open-source.
- Standard decoder-only transformer architecture chosen over mixture-of-experts for training stability.