StoryHardware
Meta details two 24,000-GPU clusters used to train Llama 3
Meta described two clusters of 24,576 NVIDIA H100 GPUs each, used to train Llama 3: one networked with RoCE Ethernet, the other with InfiniBand, to compare both at scale. Meta said it aimed to have 350,000 H100s by end of 2024, part of compute equivalent to nearly 600,000 H100s.
- Two identical 24,576-GPU clusters evaluate RoCE versus InfiniBand networking at scale.
- Uses Meta's Grand Teton open-source GPU platform and Tectonic distributed storage optimized for flash.
- Goal: 350,000 H100s by end of 2024, with total compute portfolio equivalent to nearly 600,000 H100s.
- Integrates Hammerspace parallel NFS for synchronized checkpoint saves across thousands of GPUs.