coding benchmark · included in ECI
GSO-Bench
GSO-Bench benchmark (score column: Score OPT@1).
Results
3
Random baseline
0.0%
Score ceiling
100%
Released
29 May 2025
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3 Coder 480B A35B Open source Qwen | 4.9% | — | 23 Jul 2025 | — | External ↗ | |
| 2 | Kimi K2 Instruct Open source Moonshot | 4.9% | — | 12 Jul 2025 | — | External ↗ | |
| 3 | GLM 4.5 Air Open source Z.ai | 2.9% | — | 25 Jul 2025 | — | GSO Leaderboard ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.