Skip to content
Benchmarks

7 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
DTBench
DTBench benchmark (score column: Accuracy).
reasoning205Claude Opus 5.598.9%
ARC-AGI
ARC-AGI benchmark (score column: Score).
reasoning176GPT-6 Astra98.5%
ARC-AGI-2
Abstract visual reasoning puzzles designed to be easy for humans and hard for AI.
reasoning170GPT-6 Astra95.0%
SimpleBench
Trick questions on spatio-temporal and social reasoning where humans outperform models.
reasoning98Claude Fable 581.9%
HellaSwag
HellaSwag benchmark (score column: Overall accuracy).
reasoning52GPT-4 (Mar 2023)95.3%
BBH
BBH benchmark (score column: Average).
reasoning45Gemini 1.5 Pro (May 2024)89.2%
EBR-bench
EBR-bench benchmark (score column: Best score (across scorers)).
reasoning23GPT-6 Astra76.2%