StorySoftware

xAI releases Grok 4 and Grok 4 Heavy

xAIGrok

xAI released Grok 4 and Grok 4 Heavy after reinforcement-learning training on 200,000 GPUs of its Colossus cluster. Grok 4 Heavy was the first model to score over 50% on Humanity's Last Exam; Grok 4 scored 15.9% on ARC-AGI-2. Both include native tool use, real-time search and a 256K token context window.

  • Trained with RL on 200,000 GPUs; 6x increase in compute efficiency achieved.
  • Grok 4 Heavy: first to exceed 50% on Humanity's Last Exam; 61.9% on USAMO 2025.
  • Both support native tool use, real-time search, code interpretation and web browsing.
  • Heavy variant uses parallel reasoning (multiple hypotheses simultaneously).
Read the original · Blog

More on Grok

Primer