Skip to content
coding benchmark · included in ECI

ExploitBench

ExploitBench benchmark (score column: Mean capability).
Results
3
Random baseline
0.0%
Score ceiling
100%
Released
13 May 2026

Score by model release date

Best score by organisation

#ModelScoreRelativeEvidence
1Kimi K2.6 Open source
MoonshotAI
18.4%Source ↗
2GLM 5.1 Open source
Z.ai
18.1%Source ↗
3MiniMax M2.7 Open source
MiniMax
13.3%Source ↗

Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.