StorySoftware
xAI releases Grok 4 and Grok 4 Heavy
xAI released Grok 4 and Grok 4 Heavy after reinforcement-learning training on 200,000 GPUs of its Colossus cluster. Grok 4 Heavy was the first model to score over 50% on Humanity's Last Exam; Grok 4 scored 15.9% on ARC-AGI-2. Both include native tool use, real-time search and a 256K token context window.
- Trained with RL on 200,000 GPUs; 6x increase in compute efficiency achieved.
- Grok 4 Heavy: first to exceed 50% on Humanity's Last Exam; 61.9% on USAMO 2025.
- Both support native tool use, real-time search, code interpretation and web browsing.
- Heavy variant uses parallel reasoning (multiple hypotheses simultaneously).