| Benchmark | Domain | Results | Top model | Top score | Released | Latest run | Type |
|---|---|---|---|---|---|---|---|
| GPQA diamond 198 graduate-level, 'Google-proof' multiple-choice questions in biology, physics and chemistry. | science | 262 | GPT-6 Astra | 95.8% | 20 Nov 2023 | 2 Sep 2026 | Run by Epoch |
| OTIS Mock AIME 2024-2025 Competition math problems from OTIS mock AIME exams. | math | 244 | GPT-6 Astra | 100.0% | 19 Dec 2024 | 17 Sep 2026 | Run by Epoch |
| DTBench DTBench benchmark (score column: Accuracy). | reasoning | 202 | Claude Fable 5 | 98.4% | 12 Aug 2026 | — | External leaderboard |
| Chess Puzzles Chess tactics puzzles solved by the model without tools. | games | 188 | GPT-6 Astra | 72.0% | 10 Dec 2025 | 18 Sep 2026 | Run by Epoch |
| ARC-AGI ARC-AGI benchmark (score column: Score). | reasoning | 170 | GPT-6 Astra | 98.5% | 5 Nov 2019 | — | External leaderboard |
| ARC-AGI-2 Abstract visual reasoning puzzles designed to be easy for humans and hard for AI. | reasoning | 164 | GPT-6 Astra | 95.0% | 24 Mar 2025 | — | External leaderboard |
| LMCA LMCA benchmark (score column: Score). | agents | 164 | Claude Fable 5.1 (unknown) | 65.5% | 12 Aug 2026 | — | External leaderboard |
| WeirdML Novel, unusual machine-learning tasks solved by writing and running code. | coding | 149 | GPT-6 Astra (pro, max) | 93.6% | 16 Jan 2025 | — | External leaderboard |
| MMLU MMLU benchmark (score column: EM). | knowledge | 119 | GPT-4o (Nov 2024) | 88.1% | 7 Sep 2020 | — | External leaderboard |
| Mystery Game Puzzles Mystery Game Puzzles benchmark (score column: Best score (across scorers)). | games | 107 | GPT-6 Astra | 84.0% | 30 Jul 2026 | 18 Sep 2026 | Run by Epoch |
| MATH level 5 The hardest subset of the MATH competition dataset. | math | 104 | GPT-5 | 98.1% | 5 Mar 2021 | 30 Oct 2025 | Run by Epoch |
| SimpleBench Trick questions on spatio-temporal and social reasoning where humans outperform models. | reasoning | 98 | Claude Fable 5 | 81.9% | 31 Oct 2024 | — | External leaderboard |
| FrontierMath-2025-02-28-Privatesuperseded FrontierMath-2025-02-28-Private benchmark (score column: Best score (across scorers)). | math | 97 | GPT-5.5 Pro | 52.4% | 28 Feb 2025 | 8 Jun 2026 | Run by Epoch |
| FrontierMath-Tiers-1-3-v2-Private Unpublished, expert-written research-level mathematics problems (tiers 1-3). | math | 95 | GPT-6 Astra | 93.7% | 12 Jun 2026 | 18 Sep 2026 | Run by Epoch |
| SimpleQA Verified Short factual questions testing parametric knowledge and hallucination. | knowledge | 78 | GPT-6 Astra | 75.6% | 9 Sep 2025 | 2 Sep 2026 | Run by Epoch |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded FrontierMath-Tier-4-2025-07-01-Private benchmark (score column: Best score (across scorers)). | math | 71 | AI co-mathematician | 47.9% | 11 Jul 2025 | 8 Jun 2026 | Run by Epoch |
| GSM8K GSM8K benchmark (score column: EM). | math | 68 | DeepSeek-Coder-V2 236B | 94.5% | 27 Oct 2021 | — | External leaderboard |
| ProofBench ProofBench benchmark (score column: Accuracy). | math | 68 | Claude Fable 5.1 | 100.0% | 30 Jan 2026 | — | External leaderboard |
| ARC AI2 ARC AI2 benchmark (score column: Challenge score). | other | 67 | DeepSeek V3 | 95.3% | 14 Mar 2018 | — | External leaderboard |
| Winogrande Winogrande benchmark (score column: Accuracy). | other | 64 | Llama 3.1-405B | 89.2% | 24 Jul 2019 | — | External leaderboard |
| FrontierMath-Tier-4-v2-Private The hardest FrontierMath tier: research problems that take experts days. | math | 63 | GPT-6 Astra | 97.6% | 12 Jun 2026 | 18 Sep 2026 | Run by Epoch |
| Aider polyglot Code editing exercises in C++, Go, Java, JavaScript, Python and Rust. | coding | 60 | GPT-5 | 88.0% | 21 Dec 2024 | — | External leaderboard |
| Terminal Bench Agentic tasks completed in a real terminal environment (external leaderboard). | agents | 58 | GPT-5.5 (unknown thinking) | 84.7% | 19 May 2025 | 14 May 2026 | External leaderboard |
| Fiction.LiveBench Long-context comprehension of fiction stories. | long-context | 57 | GPT-5 (medium) | 97.2% | 21 Feb 2025 | — | External leaderboard |
| DeepSWE DeepSWE benchmark (score column: Pass@1). | coding | 56 | GPT-6 Astra (xhigh) | 74.1% | 26 May 2026 | — | External leaderboard |
| HellaSwag HellaSwag benchmark (score column: Overall accuracy). | reasoning | 52 | GPT-4 (Mar 2023) | 95.3% | 19 May 2019 | — | External leaderboard |
| HLE Humanity's Last Exam: expert-written questions across many subjects. | knowledge | 48 | GPT-6 Astra (unknown thinking) | 54.8% | 23 Jan 2025 | — | External leaderboard |
| PIQA PIQA benchmark (score column: Score). | other | 47 | GPT-4o-mini | 88.7% | 26 Nov 2019 | — | External leaderboard |
| Lech Mazur Writing Lech Mazur Writing benchmark (score column: Mean score). | other | 46 | GPT-5 (medium) | 86.0% | 31 Jan 2025 | — | External leaderboard |
| METR Time Horizons Length of software tasks (in human time) an AI agent completes with 50% reliability. | agents | 46 | Claude Mythos Preview (Early) | 85.2% | 19 Mar 2025 | — | External leaderboard |
| BBH BBH benchmark (score column: Average). | reasoning | 45 | Gemini 1.5 Pro (May 2024) | 89.2% | 17 Oct 2022 | — | External leaderboard |
| Balrog Balrog benchmark (score column: Average progress). | games | 39 | GPT-6 Astra | 68.3% | 20 Nov 2024 | — | External leaderboard |
| GSO-Bench GSO-Bench benchmark (score column: Score OPT@1). | coding | 36 | Claude Opus 4.8 | 47.1% | 29 May 2025 | — | External leaderboard |
| FrontierCode FrontierCode benchmark (score column: Main score). | coding | 36 | Claude Fable 5 (unknown) | 53.5% | 8 Jun 2026 | — | External leaderboard |
| DeepResearch Bench DeepResearch Bench benchmark (score column: Average score). | agents | 35 | Claude Opus 4.6 | 55.3% | 13 Jun 2025 | — | External leaderboard |
| VPCT VPCT benchmark (score column: Correct). | multimodal | 34 | Gemini 3 Pro Preview | 91.0% | 30 Jan 2025 | — | External leaderboard |
| SWE-Bench verified 500 human-validated real GitHub issues; the model must produce a patch that passes tests. | coding | 33 | Claude Opus 4.7 | 83.5% | 13 Aug 2024 | 25 Jun 2026 | Run by Epoch |
| TriviaQA TriviaQA benchmark (score column: EM). | other | 33 | Llama 2-70B | 87.6% | 9 May 2017 | — | External leaderboard |
| APEX-Agents APEX-Agents benchmark (score column: Pass@1 score). | agents | 33 | Claude Fable 5.1 (unknown) | 68.6% | 21 Jan 2026 | — | External leaderboard |
| OpenBookQA OpenBookQA benchmark (score column: Accuracy). | other | 32 | phi-3-small 7.4B | 88.0% | 8 Sep 2018 | — | External leaderboard |
| GeoBench GeoBench benchmark (score column: ACW Country %). | multimodal | 30 | Gemini 3 Flash Preview | 88.0% | 1 Mar 2025 | — | External leaderboard |
| Surface Evolver Bench Surface Evolver Bench benchmark (score column: Mean score). | science | 27 | Kimi K3 (unknown) | 95.0% | 19 Jun 2026 | — | External leaderboard |
| Furniture Assembly Furniture Assembly benchmark (score column: Best score (across scorers)). | multimodal | 27 | Claude Opus 5.5 | 83.3% | 23 Sep 2026 | 28 Sep 2026 | Run by Epoch |
| EBR-bench EBR-bench benchmark (score column: Best score (across scorers)). | reasoning | 23 | GPT-6 Astra | 76.2% | 1 Jul 2026 | 22 Sep 2026 | Run by Epoch |
| Cybench Cybench benchmark (score column: Unguided % Solved). | coding | 22 | Claude Opus 4.6 (unknown thinking) | 93.0% | 15 Aug 2024 | — | External leaderboard |
| ScienceQA ScienceQA benchmark (score column: Score). | other | 22 | Phi-3.5-vision-instruct | 91.3% | 20 Sep 2022 | — | External leaderboard |
| CL-bench CL-bench benchmark (score column: Overall). | long-context | 22 | GPT-5.4 (xhigh) | 27.9% | 3 Feb 2026 | — | External leaderboard |
| LAMBADA LAMBADA benchmark (score column: Score). | other | 21 | Falcon-180B | 79.8% | 20 Jun 2016 | — | External leaderboard |
| CL-bench Life CL-bench Life benchmark (score column: Overall). | long-context | 16 | GPT-5.5 | 22.2% | 29 Apr 2026 | — | External leaderboard |
| Remote Labor Index Remote Labor Index benchmark (score column: Score). | agents | 15 | GPT-6 Astra (unknown thinking) | 20.8% | 29 Oct 2025 | — | External leaderboard |
| CadEval CadEval benchmark (score column: Overall pass (%)). | coding | 14 | o3 (medium) | 74.0% | 22 Apr 2025 | — | External leaderboard |
| The Agent Company The Agent Company benchmark (score column: % Resolved). | agents | 14 | DeepSeek V3.2 Exp | 42.9% | 18 Dec 2024 | — | External leaderboard |
| OSWorld 2.0 OSWorld 2.0 benchmark (score column: Binary accuracy). | agents | 13 | Claude Opus 5 | 31.4% | 26 Jun 2026 | — | External leaderboard |
| PostTrainBench PostTrainBench benchmark (score column: Average (%)). | agents | 11 | Claude Fable 5 | 41.8% | 11 Mar 2026 | — | External leaderboard |
| GDPval Economically valuable tasks across occupations, judged against expert work. | agents | 11 | GPT-5.2 (none) | 49.7% | 25 Sep 2025 | — | External leaderboard |
| ANLI ANLI benchmark (score column: Score). | other | 9 | phi-3-small 7.4B | 58.1% | 31 Oct 2019 | — | External leaderboard |
| OSWorld Computer-use tasks in real desktop operating systems. | agents | 9 | Claude Sonnet 4.6 (no thinking) | 72.1% | 11 Apr 2024 | — | External leaderboard |
| ExploitBench ExploitBench benchmark (score column: Mean capability). | coding | 9 | Claude Mythos Preview (Early) | 73.8% | 13 May 2026 | — | External leaderboard |
| MirrorCode MirrorCode benchmark (score column: Best score (across scorers)). | coding | 8 | Claude Fable 5.1 | 73.3% | 26 Jun 2026 | 10 Sep 2026 | Run by Epoch |
| SuperGLUE SuperGLUE benchmark (score column: Score). | other | 0 | — | 2 May 2019 | — | External leaderboard |