StoryTest

AMD's MI300X makes its MLPerf Inference debut on the Llama 2 70B benchmark

AMDInstinct MI300 and MI350

AMD submitted Instinct MI300X results to MLPerf Inference round v4.1 running Llama 2 70B language model with ROCm software, performing roughly in line with NVIDIA's H100 and providing buyers their first independent benchmark comparison.

  • MI300X MLPerf Inference v4.1 submission on Llama 2 70B roughly comparable to NVIDIA H100, giving customers independent performance reference.
  • AMD leveraged 192 GB HBM3 memory at 5.2 TB/s to fit full Llama 2 70B model plus KV cache without network overhead, versus H200 at 141 GB.
  • ROCm software stack optimizations including FP8 quantization retaining 99.9% accuracy and composable kernels for critical operations.
Read the original · Blog

More on Instinct MI300 and MI350

Primer