| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic | 73.3% | high | 1 Sep 2026 | 10 Sep 2026 | Source ↗ | |
| 2 | Claude Fable 5 Anthropic | 63.9% ±10.4 | high | 9 Jun 2026 | 10 Aug 2026 | Source ↗ | |
| 3 | GPT-6 Astra OpenAI | 46.7% | high | 3 Sep 2026 | 30 Aug 2026 | Source ↗ | |
| 4 | Claude Opus 4.7 Anthropic | 31.1% ±9.0 | high | 16 Apr 2026 | 12 Aug 2026 | Source ↗ | |
| 5 | GPT-5.6 Sol OpenAI | 20.0% ±9.2 | high | 9 Jul 2026 | 10 Aug 2026 | Source ↗ | |
| 6 | GPT-5.4 OpenAI | 15.6% ±7.7 | high | 5 Mar 2026 | 10 Aug 2026 | Source ↗ | |
| 7 | GPT-5.5 OpenAI | 10.0% ±6.0 | high | 23 Apr 2026 | 10 Aug 2026 | Source ↗ | |
| 8 | Gemini 3.1 Pro Preview Google | 8.9% ±4.6 | high | 19 Feb 2026 | 10 Aug 2026 | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.