agents benchmark · included in ECI
APEX-Agents
APEX-Agents benchmark (score column: Pass@1 score).
Results
9
Random baseline
0.0%
Score ceiling
100%
Released
21 Jan 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GLM-5.3-Flash (unknown) Open source Z.ai (Zhipu AI) | 52.8% | unknown | 20 Aug 2026 | — | Source ↗ | |
| 2 | Kimi K3 (unknown) Open source Moonshot | 50.6% | unknown | 16 Jul 2026 | — | Source ↗ | |
| 3 | DeepSeek V4 Pro 0813 (unknown) Open source DeepSeek | 47.3% | unknown | 13 Aug 2026 | — | Source ↗ | |
| 4 | GLM 5.1 Open source Z.ai | 40.9% | — | 7 Apr 2026 | — | Source ↗ | |
| 5 | MiniMax M3 Open source MiniMax | 37.7% | — | 1 Jun 2026 | — | Source ↗ | |
| 6 | Kimi K2.7 Code Open source MoonshotAI | 37.6% | — | 12 Jun 2026 | — | Source ↗ | |
| 7 | Inkling Open source Thinking Machines | 33.8% | — | 15 Jul 2026 | — | Source ↗ | |
| 8 | Qwen3.5 397B A17B Open source Qwen | 24.9% | — | 13 Feb 2026 | — | Source ↗ | |
| 9 | gpt-oss-120b Open source OpenAI | 4.4% | — | 5 Aug 2025 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.