StorySoftware
OpenAI previews o1, a model trained with reinforcement learning to reason before answering
OpenAI released o1-preview, a model trained to think deeply through problems before answering. It solved 83% of International Mathematics Olympiad problems versus 13% for GPT-4o, reached 89th percentile in Codeforces programming, and achieved 85.5% on MATH, 97.8% on GSM8K, and 90.8% on MMLU.
- IMO qualification exam: 83% vs. 13% for GPT-4o; achieves human expert-level reasoning
- Codeforces: 89th percentile; MATH: 85.5%; GSM8K: 97.8%; MMLU: 90.8%; GPQA Diamond: 73.3%
- Trained with RL; spends significant time thinking before answering; first in series
- Rivals human PhD-level accuracy on physics, biology, chemistry problems (GPQA)