| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.2 (none) OpenAI | 49.7% | none | 11 Dec 2025 | — | Source ↗ | |
| 2 | Claude Opus 4.5 (no thinking) Anthropic | 45.5% | — | 24 Nov 2025 | — | Source ↗ | |
| 3 | Claude Opus 4.1 Anthropic | 43.6% | — | 5 Aug 2025 | — | Source ↗ | |
| 4 | Claude Sonnet 4.5 (no thinking) Anthropic | 42.5% | — | 29 Sep 2025 | — | Source ↗ | |
| 5 | Gemini 3 Pro Preview Google DeepMind | 40.3% | — | 18 Nov 2025 | — | Source ↗ | |
| 6 | GPT-5 (medium) OpenAI | 34.8% | medium | 7 Aug 2025 | — | Source ↗ | |
| 7 | o3 (medium) OpenAI | 30.8% | medium | 16 Apr 2025 | — | Source ↗ | |
| 8 | o4 Mini OpenAI | 25.3% | high | 16 Apr 2025 | — | Source ↗ | |
| 9 | Gemini 2.5 Pro (Jun 2025) Google DeepMind | 23.3% | — | 5 Jun 2025 | — | Source ↗ | |
| 10 | Grok 4 xAI | 21.1% | high | 9 Jul 2025 | — | Source ↗ | |
| 11 | GPT-4o (Nov 2024) OpenAI | 9.9% | — | 20 Nov 2024 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.