Skip to content
Benchmarks

3 benchmarks

All scores come from the Epoch AI Benchmarking Hub (CC BY 4.0) - Epoch's own runs plus external leaderboards Epoch collects. Unrelated benchmarks are never combined into a universal score; Epoch's own composite (ECI) is shown separately on the models page with its methodology.
BenchmarkDomainResultsTop modelTop score
Fiction.LiveBench
Long-context comprehension of fiction stories.
long-context57GPT-5 (medium)97.2%
CL-bench
CL-bench benchmark (score column: Overall).
long-context22GPT-5.4 (xhigh)27.9%
CL-bench Life
CL-bench Life benchmark (score column: Overall).
long-context16GPT-5.522.2%