coding benchmark · included in ECI
SWE-Bench verified
500 human-validated real GitHub issues; the model must produce a patch that passes tests.
Results
6
Random baseline
0.0%
Score ceiling
100%
Released
13 Aug 2024
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GLM 5.2 Open source Z.ai | 78.7% ±1.9 | max | 16 Jun 2026 | 25 Jun 2026 | Log ↗ | |
| 2 | DeepSeek v4 Pro (high) Open source DeepSeek | 77.6% ±1.9 | max | 24 Apr 2026 | 18 Jun 2026 | Log ↗ | |
| 3 | Kimi K2.6 Open source MoonshotAI | 76.7% ±1.9 | — | 20 Apr 2026 | 8 May 2026 | Log ↗ | |
| 4 | GLM 5.1 Open source Z.ai | 74.2% ±2.0 | — | 7 Apr 2026 | 15 May 2026 | Source ↗ | |
| 5 | Kimi K2.5 Open source MoonshotAI | 73.8% ±2.0 | — | 27 Jan 2026 | 17 Feb 2026 | Log ↗ | |
| 6 | GLM 5 Open source Z.ai | 72.1% ±2.1 | — | 11 Feb 2026 | 15 Feb 2026 | Log ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.