| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| LMCA | agents | 41.2% | 63% | max | — | External ↗ |
| DTBench | reasoning | 90.7% | 92% | max | — | External ↗ |
| Surface Evolver Bench | science | 40.0% | 42% | high | — | External ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 45.3% ±3.0 | 48% | max | 17 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 2.4% ±2.4 | 2% | max | 17 Jun 2026 | Eval log ↗ |
| CL-bench Life | long-context | 13.5% | 61% | high | — | External ↗ |
| ProofBench | math | 16.0% | 16% | max | — | External ↗ |
| Chess Puzzles | games | 20.0% ±4.0 | 28% | max | 16 Jun 2026 | Epoch ↗ |
| SimpleQA Verified | knowledge | 47.0% ±1.6 | 62% | max | 27 Aug 2026 | Eval log ↗ |
| WeirdML | coding | 48.9% ±0.0 | 52% | max | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 96.7% ±2.0 | 97% | max | 17 Jun 2026 | Epoch ↗ |
| SWE-Bench verified | coding | 77.6% ±1.9 | 93% | max | 18 Jun 2026 | Eval log ↗ |
| GPQA diamond | science | 90.9% ±2.0 | 95% | high | 6 Aug 2026 | Eval log ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
0 observations (OpenRouter listing). History accumulates with every ingest run; a single point means no change has been observed yet.