StoryCommercial

OpenAI signs a multi-year deal for Cerebras inference capacity

CerebrasWSE-3 / CS-3

OpenAI and Cerebras agreed to deploy 750 megawatts of Cerebras wafer-scale systems for AI inference starting in 2026, in what Cerebras called the largest dedicated high-speed inference deployment in the world. The Cerebras systems deliver responses up to 15 times faster than GPU-based systems.

  • Deployment begins in 2026 across multiple phases of the multi-year agreement.
  • Cerebras systems deliver 15× faster responses than GPU-based inference on large language models.
  • OpenAI gains a dedicated low-latency inference solution to support broader AI adoption.
  • 750 megawatts is the largest single high-speed inference deployment by capacity.
Read the original · Blog

More on WSE-3 / CS-3

Primer