Benchmarks
7 benchmarks
All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
| Benchmark | Domain | Results | Top model | Top score | Released | Latest run | Type |
|---|---|---|---|---|---|---|---|
| DTBench DTBench benchmark (score column: Accuracy). | reasoning | 205 | Claude Opus 5.5 | 98.9% | 12 Aug 2026 | — | External leaderboard |
| ARC-AGI ARC-AGI benchmark (score column: Score). | reasoning | 176 | GPT-6 Astra | 98.5% | 5 Nov 2019 | — | External leaderboard |
| ARC-AGI-2 Abstract visual reasoning puzzles designed to be easy for humans and hard for AI. | reasoning | 170 | GPT-6 Astra | 95.0% | 24 Mar 2025 | — | External leaderboard |
| SimpleBench Trick questions on spatio-temporal and social reasoning where humans outperform models. | reasoning | 98 | Claude Fable 5 | 81.9% | 31 Oct 2024 | — | External leaderboard |
| HellaSwag HellaSwag benchmark (score column: Overall accuracy). | reasoning | 52 | GPT-4 (Mar 2023) | 95.3% | 19 May 2019 | — | External leaderboard |
| BBH BBH benchmark (score column: Average). | reasoning | 45 | Gemini 1.5 Pro (May 2024) | 89.2% | 17 Oct 2022 | — | External leaderboard |
| EBR-bench EBR-bench benchmark (score column: Best score (across scorers)). | reasoning | 23 | GPT-6 Astra | 76.2% | 1 Jul 2026 | 22 Sep 2026 | Run by Epoch |