AI  ·  Aug 2026  ·  Visual Essay

The Word Is
Reasoning

From 5% to 96.4% on graduate math. In 4 years. Here is the data.
5%
GPT-3 on MATH
benchmark, 2021
96.4%
o1 on MATH
benchmark, 2024
4 yrs
Time elapsed
between those two
78%
o1 on GPQA Diamond
PhD-level science

Where Every Model Stands Right Now

Five frontier models across four reasoning benchmarks. Hover any cell to see the exact score and its source.

Benchmark Heatmap: Score (%) by Model and Task
Sources: Hendrycks et al. arXiv:2103.03874 (MATH); Hendrycks et al. arXiv:2009.03300 (MMLU); Chen et al. arXiv:2107.03374 (HumanEval); Rein et al. arXiv:2311.12022 (GPQA Diamond). Model cards: GPT-4 (arXiv:2303.08774), Claude 3.5 Sonnet (Anthropic 2024), Gemini 1.5 Pro (arXiv:2403.05530), o1 (OpenAI system card Sep 2024), DeepSeek R1 (arXiv:2501.12948).

The MATH Climb

One benchmark. Four years. The steepest capability jump in benchmarking history. Hover any point for the score and date.

MATH Benchmark Score (%) Over Time
GPT-3 (2021): Hendrycks et al. arXiv:2103.03874. GPT-4 (2023): arXiv:2303.08774. Claude 3.5 Sonnet (2024): Anthropic model card. o1 (late 2024): OpenAI o1 system card.

Math vs. Science: Who Generalizes?

A model that reasons well on math should reason well on PhD-level science. The scatter confirms it. Hover any dot.

MATH Score (x-axis) vs. GPQA Diamond Score (y-axis)
GPQA Diamond tests PhD-level biology, chemistry, and physics. Correlation between MATH and GPQA scores suggests reasoning transfers across domains. Sources as above.

Before and After Reasoning Models

GPT-4 vs. o1. Same company. One architectural shift. The gap is not incremental. Hover any bar.

GPT-4 vs. o1: Benchmark Comparison (%)
GPT-4 scores from arXiv:2303.08774. o1 scores from OpenAI o1 system card, September 2024.

The Reasoning Timeline

Six inflection points that changed what AI can think through. Each bar represents time elapsed since the previous event.

Key Milestones in AI Reasoning Capability (2021-2025)
Chain-of-Thought prompting: Wei et al. arXiv:2201.11903 (NeurIPS 2022). GPT-4: arXiv:2303.08774. DeepSeek R1: arXiv:2501.12948.
19x
MATH score gain
2021 to 2024
97.3%
DeepSeek R1
on MATH, Jan 2025
Chain-of-Thought
The technique that
started it all, 2022

Every number is cited.

Want to apply reasoning AI in your enterprise?

Arjun Jaggi advises C-suite teams on AI infrastructure, model selection, and enterprise AI strategy.

Schedule a conversation