StoryCommercial
SambaNova Cloud launches, hosting Meta's Llama models at 655 tokens/sec for Llama 4 Maverick
SambaNova launched SambaCloud, a cloud platform running the RDU accelerator architecture, hosting Meta's Llama language models. Llama 4 Maverick achieves 655 tokens per second, Llama 3.1 405B at 132 tokens/sec, and Llama 3.1 70B at 461 tokens/sec, demonstrating throughput competitive with or exceeding GPU-based inference.
- Llama 4 Maverick: 655 tokens/sec
- Llama 3.1 405B: 132 tokens/sec
- Llama 3.1 70B: 461 tokens/sec
- Demonstrates dataflow advantage for inference throughput