StoryCommercial

Cerebras launches an inference service claiming record tokens-per-second speeds

CerebrasWSE-3 / CS-3

Cerebras launched an inference service running Llama 3.1 models at 1,800 tokens per second for the 8B parameter model and 450 for the 70B, claiming 20 times faster performance than GPU-based clouds. Keeping model weights in the wafer-scale chip's on-chip memory eliminates the memory bottleneck that limits GPUs.

  • Llama 3.1 8B inference at 1,800 tokens/second and 70B at 450 tokens/second; claimed 20x faster than NVIDIA GPU clouds and 2.4x faster than Groq on 8B.
  • Wafer Scale Engine 3 uses on-chip memory for full model weights, avoiding network traffic and memory bottlenecks of GPU-based inference.
  • Native 16-bit weights maintain superior accuracy versus 8-bit quantized alternatives while delivering record throughput per token.
Read the original · Blog

More on WSE-3 / CS-3

Primer