coding benchmark · included in ECI
ExploitBench
ExploitBench benchmark (score column: Mean capability).
Results
3
Random baseline
0.0%
Score ceiling
100%
Released
13 May 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K2.6 Open source MoonshotAI | 18.4% | — | 20 Apr 2026 | — | Source ↗ | |
| 2 | GLM 5.1 Open source Z.ai | 18.1% | — | 7 Apr 2026 | — | Source ↗ | |
| 3 | MiniMax M2.7 Open source MiniMax | 13.3% | — | 18 Mar 2026 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.