math benchmark · included in ECI
FrontierMath-Tier-4-v2-Private
The hardest FrontierMath tier: research problems that take experts days.
Results
11
Random baseline
0.0%
Score ceiling
100%
Released
12 Jun 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K3 Open source MoonshotAI | 39.0% ±7.7 | max | 16 Jul 2026 | 17 Jul 2026 | Log ↗ | |
| 2 | GLM 5.3 Open source Z.ai | 29.3% ±7.2 | max | 14 Aug 2026 | 25 Aug 2026 | Source ↗ | |
| 3 | GLM 5.2 Open source Z.ai | 29.3% ±7.2 | max | 16 Jun 2026 | 19 Jun 2026 | Log ↗ | |
| 4 | DeepSeek V4 Pro 0813 Open source DeepSeek | 26.8% ±7.0 | max | 13 Aug 2026 | 19 Aug 2026 | Log ↗ | |
| 5 | Kimi K2.6 Open source MoonshotAI | 25.6% ±7.1 | — | 20 Apr 2026 | 10 Jun 2026 | Log ↗ | |
| 6 | DeepSeek V4 Flash 0731 Open source DeepSeek | 24.4% ±6.8 | max | 31 Jul 2026 | 2 Aug 2026 | Log ↗ | |
| 7 | GLM 5.3 Flash Open source Z.ai | 17.1% ±5.9 | max | 20 Aug 2026 | 27 Aug 2026 | Source ↗ | |
| 8 | Inkling Small (xhigh) Open source Thinking Machines | 17.1% ±5.9 | xhigh | 15 Jul 2026 | 14 Aug 2026 | Log ↗ | |
| 9 | Kimi K2.7 Code Open source MoonshotAI | 12.2% ±5.2 | — | 12 Jun 2026 | 13 Jun 2026 | Log ↗ | |
| 10 | Inkling (xhigh) Open source Thinking Machines | 4.9% ±3.4 | xhigh | 15 Jul 2026 | 6 Aug 2026 | Log ↗ | |
| 11 | DeepSeek v4 Pro (high) Open source DeepSeek | 2.4% ±2.4 | max | 24 Apr 2026 | 17 Jun 2026 | Log ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.