openai-o1-83pct-imo-vs-gpt4o-13pct

IN premiseentries/2026/06/21/wiki-Large_language_model-chunk-2.md

Created 2026-06-21T09:50:09+00:00

OpenAI o1 scored 83% on IMO qualifying problems compared to GPT-4o's 13%.

Summary

OpenAI's o1 model solved 83% of International Mathematical Olympiad qualifying problems while GPT-4o managed only 13%, a gap that points to a fundamentally different and far more effective approach to deep, multi-step mathematical reasoning in o1. This observation serves as a hard data point for any claim about where the frontier of AI math capability actually sits and how large a generational leap in reasoning occurred between these two models.

Dependents

These beliefs depend on this one: