StoryCommercial
OpenAI signs a multi-year deal for Cerebras inference capacity
OpenAI and Cerebras agreed to deploy 750 megawatts of Cerebras wafer-scale systems for AI inference starting in 2026, in what Cerebras called the largest dedicated high-speed inference deployment in the world. The Cerebras systems deliver responses up to 15 times faster than GPU-based systems.
- Deployment begins in 2026 across multiple phases of the multi-year agreement.
- Cerebras systems deliver 15× faster responses than GPU-based inference on large language models.
- OpenAI gains a dedicated low-latency inference solution to support broader AI adoption.
- 750 megawatts is the largest single high-speed inference deployment by capacity.