AI benchmarks
AI benchmarks are standard test sets used to measure model capabilities such as reasoning, maths or coding, so models can be compared on the same tasks.
Each benchmark measures something narrow: GPQA Diamond tests graduate-level science questions, AIME competition maths, DeepSWE real software-engineering tasks. No single benchmark says which model is "best" overall, and results depend on settings.
Composite measures such as Epoch AI's Capabilities Index combine many benchmarks statistically. This site shows each measure separately with its source.
Related ranking: LLM leaderboard
- 1.GPT-6 Astra OpenAI166.6
- 2.Claude Fable 5.1 Anthropic165.0
- 3.Claude Fable 5 Anthropic163.6
- 4.Claude Opus 5 Anthropic162.7
- 5.GPT-5.5 Pro OpenAI162.4
Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Full ranking
Explanation written by AI Stats Live editors; last reviewed 29 Sep 2026. Examples and rankings update automatically from the sources listed on the methodology page.