StorySoftware
Meta releases Llama 3 in 8B and 70B sizes
Meta released Llama 3 in 8B and 70B parameter sizes, trained on over 15 trillion tokens—seven times Llama 2's data—with an 8,192-token context window. A 400B+ parameter model was still in training; Meta promised larger, multimodal and longer-context versions within months.
- Trained on 15 trillion tokens from public sources, including four times more code than Llama 2.
- Includes support for over 30 languages; tokenizer expanded to 128K vocabulary for efficiency.
- Implements grouped query attention for improved inference efficiency.
- 405B+ model in development with planned multimodality, multi-language fluency, longer context.