openai-o1-83pct-imo-vs-gpt4o-13pct
IN premise — entries/2026/06/21/wiki-Large_language_model-chunk-2.md
Created 2026-06-21T09:50:09+00:00
OpenAI o1 scored 83% on IMO qualifying problems compared to GPT-4o's 13%.
Summary
OpenAI's o1 model solved 83% of International Mathematical Olympiad qualifying problems while GPT-4o managed only 13%, a gap that points to a fundamentally different and far more effective approach to deep, multi-step mathematical reasoning in o1. This observation serves as a hard data point for any claim about where the frontier of AI math capability actually sits and how large a generational leap in reasoning occurred between these two models.
Dependents
These beliefs depend on this one:
- IN reasoning-models-represent-distinct-capability-tier — Reasoning-specialized models — OpenAI o1 scoring 83% vs GPT-4o's 13% on IMO qualifying problems, DeepSeek R1 matching proprietary models at lower cost — represent a distinct capability tier above standard LLMs, achievable through both proprietary and open-weight approaches.