| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 Anthropic | 55.3% | high | 5 Feb 2026 | — | Source ↗ | |
| 2 | Claude Sonnet 4.6 Anthropic | 54.9% | high | 17 Feb 2026 | — | Source ↗ | |
| 3 | Claude Opus 4.5 Anthropic | 54.8% | high | 24 Nov 2025 | — | Source ↗ | |
| 4 | GPT-5.5 OpenAI | 54.0% | high | 23 Apr 2026 | — | Source ↗ | |
| 5 | Claude Opus 4.5 (low) Anthropic | 53.7% | low | 24 Nov 2025 | — | Source ↗ | |
| 6 | Claude Opus 4.6 (medium) Anthropic | 53.2% | medium | 5 Feb 2026 | — | Source ↗ | |
| 7 | Claude Sonnet 4.5 (2k thinking) Anthropic | 52.6% | 2K | 29 Sep 2025 | — | Source ↗ | |
| 8 | Claude Opus 4.6 (low) Anthropic | 51.4% | low | 5 Feb 2026 | — | Source ↗ | |
| 9 | Claude Sonnet 4.6 (low) Anthropic | 50.4% | low | 17 Feb 2026 | — | Source ↗ | |
| 10 | Claude Opus 4.8 Anthropic | 50.2% | high | 28 May 2026 | — | Source ↗ | |
| 11 | Gemini 3 Flash Preview Google | 49.8% | low | 17 Dec 2025 | — | Source ↗ | |
| 12 | GPT-5.5 (medium) OpenAI | 49.6% | medium | 23 Apr 2026 | — | Source ↗ | |
| 13 | GPT-5 (low) OpenAI | 49.6% | low | 7 Aug 2025 | — | Source ↗ | |
| 14 | GPT-5 (minimal) OpenAI | 48.9% | minimal | 7 Aug 2025 | — | Source ↗ | |
| 15 | GPT-5.5 (low) OpenAI | 48.7% | low | 23 Apr 2026 | — | Source ↗ | |
| 16 | GPT-5 (medium) OpenAI | 48.6% | medium | 7 Aug 2025 | — | Source ↗ | |
| 17 | Claude Opus 4.1 Anthropic | 48.3% | — | 5 Aug 2025 | — | Source ↗ | |
| 18 | GPT-5 OpenAI | 48.1% | high | 7 Aug 2025 | — | Source ↗ | |
| 19 | Gemini 3.1 Pro Preview Google | 47.8% | high | 19 Feb 2026 | — | Source ↗ | |
| 20 | Claude Sonnet 4.5 (low) Anthropic | 47.5% | low | 29 Sep 2025 | — | Source ↗ | |
| 21 | Grok 4 xAI | 47.3% | agent: True | 9 Jul 2025 | — | Source ↗ | |
| 22 | Claude Opus 4 Anthropic | 46.8% | 2K | 22 May 2025 | — | Source ↗ | |
| 23 | Claude Sonnet 4 Anthropic | 46.6% | 2K | 22 May 2025 | — | Source ↗ | |
| 24 | Gemini 3 Pro Preview (low) Google DeepMind | 46.3% | low | 18 Nov 2025 | — | Source ↗ | |
| 25 | Claude Haiku 4.5 (low) Anthropic | 45.5% | low | 15 Oct 2025 | — | Source ↗ | |
| 26 | o3 (medium) OpenAI | 45.2% | medium | 16 Apr 2025 | — | Source ↗ | |
| 27 | Claude 3.7 Sonnet (2k thinking) Anthropic | 43.6% | 2K | 24 Feb 2025 | — | Source ↗ | |
| 28 | GPT-5.1 (low) OpenAI | 42.8% | low | 13 Nov 2025 | — | Source ↗ | |
| 29 | Gemini 2.5 Pro Preview (Jun 2025) Google DeepMind | 42.8% | — | 5 Jun 2025 | — | Source ↗ | |
| 30 | Gemini 2.5 Pro (Jun 2025) Google DeepMind | 41.5% | — | 5 Jun 2025 | — | Source ↗ | |
| 31 | GPT-5.2 (low) OpenAI | 41.1% | low | 11 Dec 2025 | — | Source ↗ | |
| 32 | Gemini 3.1 Flash Lite Google | 37.3% | low | 3 Mar 2026 | — | Source ↗ | |
| 33 | GPT-5.4 mini (low) OpenAI | 36.3% | low | 17 Mar 2026 | — | Source ↗ | |
| 34 | GPT-5.4 (low) OpenAI | 35.1% | low | 5 Mar 2026 | — | Source ↗ | |
| 35 | DeepSeek-R1 (May 2025) Open source DeepSeek | 35.1% | — | 28 May 2025 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.