agents benchmark · included in ECI
APEX-Agents
APEX-Agents benchmark (score column: Pass@1 score).
Results
34
Random baseline
0.0%
Score ceiling
100%
Released
21 Jan 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 73.5% | max | 22 Sep 2026 | — | Source ↗ | |
| 2 | Claude Fable 5.1 (unknown) Anthropic | 68.6% | unknown | 1 Sep 2026 | — | Source ↗ | |
| 3 | Gemini 3.7 Flash (unknown) Google DeepMind | 67.8% | unknown | 13 Aug 2026 | — | Source ↗ | |
| 4 | Claude Opus 5 Anthropic | 65.8% | max | 24 Jul 2026 | — | Source ↗ | |
| 5 | Grok 4.6 (unknown) xAI | 65.3% | unknown | 12 Aug 2026 | — | Source ↗ | |
| 6 | GPT-6 Astra (unknown thinking) OpenAI | 64.7% | unknown | 3 Sep 2026 | — | Source ↗ | |
| 7 | Gemini 3.8 Flash (unknown) Google DeepMind | 64.3% | unknown | 2 Sep 2026 | — | Source ↗ | |
| 8 | Claude Fable 5 Anthropic | 63.6% | — | 9 Jun 2026 | — | Source ↗ | |
| 9 | Claude Fable 5.1 Anthropic | 59.7% | high | 1 Sep 2026 | — | Source ↗ | |
| 10 | GPT-5.6 Terra OpenAI | 58.2% | max | 9 Jul 2026 | — | Source ↗ | |
| 11 | Muse Spark 1.3 (unknown) Meta AI | 57.8% | unknown | 2 Sep 2026 | — | Source ↗ | |
| 12 | GLM-5.3 (unknown thinking) Z.ai (Zhipu AI) | 56.6% | unknown | 14 Aug 2026 | — | Source ↗ | |
| 13 | Grok 4.5 (unknown thinking) xAI | 56.2% | unknown | 8 Jul 2026 | — | Source ↗ | |
| 14 | GPT-5.5 (unknown thinking) OpenAI | 55.1% | unknown | 23 Apr 2026 | — | Source ↗ | |
| 15 | Claude Sonnet 5 (unknown thinking) Anthropic | 54.5% | unknown | 30 Jun 2026 | — | Source ↗ | |
| 16 | GLM-5.3-Flash (unknown) Open source Z.ai (Zhipu AI) | 52.8% | unknown | 20 Aug 2026 | — | Source ↗ | |
| 17 | GPT-5.4 (unknown thinking) OpenAI | 52.4% | unknown | 5 Mar 2026 | — | Source ↗ | |
| 18 | GPT-5.6 Sol (pro, max) OpenAI | 51.4% | promax | 9 Jul 2026 | — | Source ↗ | |
| 19 | Kimi K3 (unknown) Open source Moonshot | 50.6% | unknown | 16 Jul 2026 | — | Source ↗ | |
| 20 | Claude Opus 4.7 Anthropic | 49.2% | max | 16 Apr 2026 | — | Source ↗ | |
| 21 | Claude Opus 4.8 Anthropic | 48.9% | max | 28 May 2026 | — | Source ↗ | |
| 22 | DeepSeek V4 Pro 0813 (unknown) Open source DeepSeek | 47.3% | unknown | 13 Aug 2026 | — | Source ↗ | |
| 23 | Gemini 3.6 Flash Google | 46.9% | unknown | 21 Jul 2026 | — | Source ↗ | |
| 24 | Claude Opus 4.6 Anthropic | 46.3% | max | 5 Feb 2026 | — | Source ↗ | |
| 25 | Claude Sonnet 4.6 Anthropic | 43.0% | high | 17 Feb 2026 | — | Source ↗ | |
| 26 | GLM 5.1 Open source Z.ai | 40.9% | — | 7 Apr 2026 | — | Source ↗ | |
| 27 | MiniMax M3 Open source MiniMax | 37.7% | — | 1 Jun 2026 | — | Source ↗ | |
| 28 | Kimi K2.7 Code Open source MoonshotAI | 37.6% | — | 12 Jun 2026 | — | Source ↗ | |
| 29 | Muse Spark 1.2 Meta | 36.4% | — | 5 Aug 2026 | — | Source ↗ | |
| 30 | Gemini 3.1 Pro Preview Google | 35.3% | — | 19 Feb 2026 | — | Source ↗ | |
| 31 | Inkling Open source Thinking Machines | 33.8% | — | 15 Jul 2026 | — | Source ↗ | |
| 32 | Gemini 3.5 Flash Google | 27.5% | unknown | 19 May 2026 | — | Source ↗ | |
| 33 | Qwen3.5 397B A17B Open source Qwen | 24.9% | — | 13 Feb 2026 | — | Source ↗ | |
| 34 | gpt-oss-120b Open source OpenAI | 4.4% | — | 5 Aug 2025 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.