Skip to content
Benchmarks

11 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
LMCA
LMCA benchmark (score column: Score).
agents167Claude Opus 5.568.2%
Terminal Bench
Agentic tasks completed in a real terminal environment (external leaderboard).
agents58GPT-5.5 (unknown thinking)84.7%
METR Time Horizons
Length of software tasks (in human time) an AI agent completes with 50% reliability.
agents46Claude Mythos Preview (Early)85.2%
APEX-Agents
APEX-Agents benchmark (score column: Pass@1 score).
agents38Claude Sonnet 5.575.5%
DeepResearch Bench
DeepResearch Bench benchmark (score column: Average score).
agents35Claude Opus 4.655.3%
Remote Labor Index
Remote Labor Index benchmark (score column: Score).
agents15GPT-6 Astra (unknown thinking)20.8%
The Agent Company
The Agent Company benchmark (score column: % Resolved).
agents14DeepSeek V3.2 Exp42.9%
OSWorld 2.0
OSWorld 2.0 benchmark (score column: Binary accuracy).
agents13Claude Opus 531.4%
PostTrainBench
PostTrainBench benchmark (score column: Average (%)).
agents11Claude Fable 541.8%
GDPval
Economically valuable tasks across occupations, judged against expert work.
agents11GPT-5.2 (none)49.7%
OSWorld
Computer-use tasks in real desktop operating systems.
agents9Claude Sonnet 4.6 (no thinking)72.1%