The brief
- Shifted focus in 2025 from training to AI inference with AI200 and AI250 accelerators for data centers.
- HUMAIN, Saudi Arabia's public AI company, is first customer, planning 200 MW of Qualcomm racks starting 2026.
- Targets $15B in data-center revenue by 2029; partnerships with Meta for CPUs and Microsoft for networking.
Technical approach
LPDDR memory for token generation
AI200 holds 768 GB of LPDDR5X per card, optimized for inference where memory bandwidth limits throughput. 160 kW per rack.
Purpose-built for LLMs and multimodal models
Unlike training chips, AI200 and AI250 rack systems prioritize generating tokens efficiently. Qualcomm targets the least power per token, not peak TFLOPS.
Liquid cooling with PCIe and Ethernet scaling
Direct liquid cooling manages 160 kW per rack. Cards scale within a rack via PCIe, racks link across a data center via Ethernet.


