math benchmark · included in ECI
GSM8K
GSM8K benchmark (score column: EM).
Results
68
Random baseline
0.0%
Score ceiling
100%
Released
27 Oct 2021
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek-Coder-V2 236B Open source DeepSeek | 94.5% | — | 17 Jun 2024 | — | Source ↗ | |
| 2 | Qwen2.5-Coder-14B-Instruct | 94.2% | — | 6 Nov 2024 | — | Source ↗ | |
| 3 | Qwen2.5 Coder 32B Instruct Open source qwen | 93.0% | — | 11 Nov 2024 | — | Source ↗ | |
| 4 | GPT-4 (Mar 2023) OpenAI | 92.0% | — | 14 Mar 2023 | — | Source ↗ | |
| 5 | GPT-4o-mini OpenAI | 91.3% | — | 18 Jul 2024 | — | Source ↗ | |
| 6 | Qwen2.5-Coder-32B Open source Alibaba | 91.1% | — | 18 Sep 2024 | — | Source ↗ | |
| 7 | GPT-4 (Jun 2023) OpenAI | 90.0% | — | 13 Jun 2023 | — | Source ↗ | |
| 8 | Qwen2.5-Coder-14B | 88.7% | — | 18 Sep 2024 | — | Source ↗ | |
| 9 | Phi-3.5-MoE Open source Microsoft | 88.7% | — | 17 Aug 2024 | — | Source ↗ | |
| 10 | DeepSeek-Coder-V2-Lite-Instruct | 87.6% | — | 13 Jun 2024 | — | Source ↗ | |
| 11 | Qwen2.5-Coder (7B) Open source Alibaba | 86.7% | — | 18 Sep 2024 | — | Source ↗ | |
| 12 | Claude Instant Anthropic | 86.7% | — | 9 Aug 2023 | — | Source ↗ | |
| 13 | Phi-3.5-mini Open source Microsoft | 86.2% | — | 16 Aug 2024 | — | Source ↗ | |
| 14 | Gemma 2 9B Open source Google DeepMind | 84.9% | — | 24 Jun 2024 | — | Source ↗ | |
| 15 | Mistral Nemo Open source Mistral | 84.2% | — | 18 Jul 2024 | — | Source ↗ | |
| 16 | Llama 3.1-8B Open source Meta AI | 82.4% | — | 23 Jul 2024 | — | Source ↗ | |
| 17 | Gemini 1.5 Flash (May 2024) Google DeepMind | 82.4% | — | 23 May 2024 | — | Source ↗ | |
| 18 | Qwen2.5-Coder-3B-Instruct | 80.7% | — | 6 Nov 2024 | — | Source ↗ | |
| 19 | Yi-34B Open source 01.AI | 76.0% | — | 2 Nov 2023 | — | Source ↗ | |
| 20 | Qwen2.5-Coder-3B | 75.7% | — | 18 Sep 2024 | — | Source ↗ | |
| 21 | Mixtral 8x7B Open source Mistral AI | 74.4% | — | 11 Dec 2023 | — | Source ↗ | |
| 22 | Stable Beluga 2 Open source Stability AI | 69.6% | — | 20 Jul 2023 | — | Source ↗ | |
| 23 | Llama 2-70B Open source Meta AI | 69.6% | — | 18 Jul 2023 | — | Source ↗ | |
| 24 | DeepSeek-Coder-V2-Lite-Base | 67.1% | — | 13 Jun 2024 | — | Source ↗ | |
| 25 | Qwen2.5-Coder (1.5B) Open source Alibaba | 65.8% | — | 18 Sep 2024 | — | Source ↗ | |
| 26 | internlm-20b | 62.9% | — | 18 Sep 2023 | — | Source ↗ | |
| 27 | Qwen-14B Open source Alibaba | 61.3% | — | 24 Sep 2023 | — | Source ↗ | |
| 28 | GPT-3.5 Turbo (June 2023) OpenAI | 57.8% | — | 13 Jun 2023 | — | Source ↗ | |
| 29 | StarCoder 2 15B Open source Hugging Face,ServiceNow,NVIDIA,BigCode | 57.7% | — | 20 Feb 2024 | — | Source ↗ | |
| 30 | Mistral 7B v0.1 Open source Mistral AI | 54.4% | — | 27 Sep 2023 | — | Source ↗ | |
| 31 | Falcon-180B Open source Technology Innovation Institute | 54.4% | — | 6 Sep 2023 | — | Source ↗ | |
| 32 | LLaMA-65B Open source Meta AI | 54.4% | — | 24 Feb 2023 | — | Source ↗ | |
| 33 | Falcon 2 11B Open source Technology Innovation Institute | 53.8% | — | 9 May 2024 | — | Source ↗ | |
| 34 | Baichuan2-13B Open source Baichuan | 52.8% | — | 6 Sep 2023 | — | Source ↗ | |
| 35 | Qwen-7B Open source Alibaba | 51.7% | — | 28 Sep 2023 | — | Source ↗ | |
| 36 | Gemma 7B Open source Google DeepMind | 46.4% | — | 21 Feb 2024 | — | Source ↗ | |
| 37 | Nemotron-4 15B NVIDIA | 46.0% | — | 26 Feb 2024 | — | Source ↗ | |
| 38 | Baichuan2-13B-Chat | 45.7% | — | 6 Sep 2023 | — | Source ↗ | |
| 39 | Yi 6B Open source 01.AI | 44.9% | — | 22 Nov 2023 | — | Source ↗ | |
| 40 | LLaMA-33B Open source Meta AI | 44.1% | — | 24 Feb 2023 | — | Source ↗ | |
| 41 | Llama 2-34B Meta AI | 42.2% | — | 18 Jul 2023 | — | Source ↗ | |
| 42 | INTELLECT-1 Prime Intellect,Hugging Face,Arcee AI | 38.6% | — | 29 Nov 2024 | — | Source ↗ | |
| 43 | CodeQwen1.5-7B | 37.7% | — | 15 Apr 2024 | — | Source ↗ | |
| 44 | Llama 2-13B Open source Meta AI | 36.9% | — | 18 Jul 2023 | — | Source ↗ | |
| 45 | Mistral 7B v0.2 Open source Mistral AI | 35.4% | — | 11 Dec 2023 | — | Source ↗ | |
| 46 | DeepSeek Coder 33B Open source DeepSeek,Peking University | 35.4% | — | 2 Nov 2023 | — | Source ↗ | |
| 47 | Qwen2.5-Coder-0.5B | 34.5% | — | 18 Sep 2024 | — | Source ↗ | |
| 48 | MPT-30B Open source MosaicML | 34.4% | — | 22 Jun 2023 | — | Source ↗ | |
| 49 | Falcon-40B Open source Technology Innovation Institute | 33.8% | — | 25 May 2023 | — | Source ↗ | |
| 50 | StarCoder 2 7B Open source Hugging Face,ServiceNow,NVIDIA,BigCode | 32.7% | — | 20 Feb 2024 | — | Source ↗ | |
| 51 | chatglm2-6b | 32.4% | — | 24 Jun 2023 | — | Source ↗ | |
| 52 | internlm-7b | 31.2% | — | 5 Jul 2023 | — | Source ↗ | |
| 53 | vicuna-13b-v1.1 | 28.1% | — | 12 Apr 2023 | — | Source ↗ | |
| 54 | Baichuan 1-13B Open source Baichuan | 26.8% | — | 11 Jul 2023 | — | Source ↗ | |
| 55 | Baichuan 2-7B Open source Baichuan | 24.6% | — | 20 Sep 2023 | — | Source ↗ | |
| 56 | Vicuna-13B-v1.3 Open source Large Model Systems Organization,University of California (UC) Berkeley | 22.6% | — | 18 Jun 2023 | — | Source ↗ | |
| 57 | StarCoder 2 3B Open source Hugging Face,ServiceNow,NVIDIA,BigCode | 21.6% | — | 22 Feb 2024 | — | Source ↗ | |
| 58 | DeepSeek Coder 6.7B Open source DeepSeek,Peking University | 21.3% | — | 2 Nov 2023 | — | Source ↗ | |
| 59 | Qwen-1_8B | 21.2% | 8B | 30 Nov 2023 | — | Source ↗ | |
| 60 | LLaMA-13B Open source Meta AI | 20.5% | — | 24 Feb 2023 | — | Source ↗ | |
| 61 | Gemma 2B Open source Google DeepMind | 17.7% | — | 21 Feb 2024 | — | Source ↗ | |
| 62 | Llama 2-7B Open source Meta AI | 16.7% | — | 18 Jul 2023 | — | Source ↗ | |
| 63 | internlm-chat-20b | 15.7% | — | 17 Sep 2023 | — | Source ↗ | |
| 64 | LLaMA-7B Open source Meta AI | 11.0% | — | 24 Feb 2023 | — | Source ↗ | |
| 65 | Baichuan1-7B Open source Baichuan | 9.2% | — | 1 Jun 2023 | — | Source ↗ | |
| 66 | MPT-7B Open source MosaicML | 9.1% | — | 5 May 2023 | — | Source ↗ | |
| 67 | Falcon-7B Open source Technology Innovation Institute | 6.8% | — | 24 Apr 2023 | — | Source ↗ | |
| 68 | DeepSeek Coder 1.3B Open source DeepSeek,Peking University | 4.4% | — | 2 Nov 2023 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.