StoryCommercial

SambaNova Cloud launches, hosting Meta's Llama models at 655 tokens/sec for Llama 4 Maverick

SambaNova

SambaNova launched SambaCloud, a cloud platform running the RDU accelerator architecture, hosting Meta's Llama language models. Llama 4 Maverick achieves 655 tokens per second, Llama 3.1 405B at 132 tokens/sec, and Llama 3.1 70B at 461 tokens/sec, demonstrating throughput competitive with or exceeding GPU-based inference.

  • Llama 4 Maverick: 655 tokens/sec
  • Llama 3.1 405B: 132 tokens/sec
  • Llama 3.1 70B: 461 tokens/sec
  • Demonstrates dataflow advantage for inference throughput
Read the original · Press

More on SambaNova

Primer