coding benchmark · included in ECI
Aider polyglot
Code editing exercises in C++, Go, Java, JavaScript, Python and Rust.
Results
19
Random baseline
0.0%
Score ceiling
100%
Released
21 Dec 2024
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V3.2 Open source DeepSeek | 74.2% | — | 1 Dec 2025 | — | External ↗ | |
| 2 | DeepSeek V3.2 Exp Open source DeepSeek | 74.2% | thinking | 29 Sep 2025 | — | Aider LLM Leaderboards ↗ | |
| 3 | DeepSeek-R1 (May 2025) Open source DeepSeek | 71.4% | — | 28 May 2025 | — | External ↗ | |
| 4 | Qwen3-235B-A22B-Instruct (Jul 2025) Open source Alibaba | 59.6% | — | 25 Jul 2025 | — | External ↗ | |
| 5 | Qwen3 235B A22B Open source Qwen | 59.6% | — | 28 Apr 2025 | — | Aider LLM Leaderboards ↗ | |
| 6 | Kimi K2 0905 (Novita) Open source Moonshot | 59.1% | — | 5 Sep 2025 | — | External ↗ | |
| 7 | Kimi K2 Instruct Open source Moonshot | 59.1% | — | 12 Jul 2025 | — | Aider LLM Leaderboards ↗ | |
| 8 | R1 Open source DeepSeek | 56.9% | — | 20 Jan 2025 | — | External ↗ | |
| 9 | DeepSeek-V3 (Mar 2025) Open source DeepSeek | 55.1% | — | 24 Mar 2025 | — | External ↗ | |
| 10 | DeepSeek V3 Open source DeepSeek | 48.4% | — | 26 Dec 2024 | — | Aider LLM Leaderboards ↗ | |
| 11 | gpt-oss-120b Open source OpenAI | 41.8% | high | 5 Aug 2025 | — | Aider LLM Leaderboards ↗ | |
| 12 | Qwen3 32B Open source Qwen | 40.0% | — | 29 Apr 2025 | — | External ↗ | |
| 13 | QwQ-32B Open source Alibaba | 20.9% | — | 5 Mar 2025 | — | External ↗ | |
| 14 | DeepSeek-V2.5 (Sep 2024) Open source DeepSeek | 17.8% | — | 6 Sep 2024 | — | External ↗ | |
| 15 | Qwen2.5 Coder 32B Instruct Open source qwen | 16.4% | — | 11 Nov 2024 | — | Aider LLM Leaderboards ↗ | |
| 16 | Llama 4 Maverick Open source Meta | 15.6% | — | 6 Apr 2025 | — | External ↗ | |
| 17 | Cohere Command A Open source Cohere | 12.0% | — | 13 Mar 2025 | — | External ↗ | |
| 18 | Codestral Open source Mistral AI | 11.1% | — | 13 Jan 2025 | — | External ↗ | |
| 19 | Gemma 3 27B Open source Google | 4.9% | — | 12 Mar 2025 | — | External ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.