StoryTest
AMD's MI300X makes its MLPerf Inference debut on the Llama 2 70B benchmark
AMD submitted Instinct MI300X results to MLPerf Inference round v4.1 running Llama 2 70B language model with ROCm software, performing roughly in line with NVIDIA's H100 and providing buyers their first independent benchmark comparison.
- MI300X MLPerf Inference v4.1 submission on Llama 2 70B roughly comparable to NVIDIA H100, giving customers independent performance reference.
- AMD leveraged 192 GB HBM3 memory at 5.2 TB/s to fit full Llama 2 70B model plus KV cache without network overhead, versus H200 at 141 GB.
- ROCm software stack optimizations including FP8 quantization retaining 99.9% accuracy and composable kernels for critical operations.