Skip to content

Best for math

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics).

Ranked by Epoch's mock AIME 2024-2025 (competition mathematics). Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101o3-minimediumOpenAI—74.3———
102Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
103Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
104Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
105QwQ-32BAlibaba (Qwen)137.665.3———
106GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
107Claude 3.7 Sonnet64k thinkingAnthropic · +3 variants—79.7———
108Gemini 2.0 Flash Thinking ExpGoogle—57.1———
109Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
110Qwen3.5-4BAlibaba (Qwen)—————
111Grok 3xAI138.375.8———
112DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
113R1DeepSeek139.071.7—64K$1.15
114qwen3-4b-instruct-2507—45.8———
115DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
116DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
117o1-miniOpenAI · +1 variant135.862.4———
118GPT-4.1 MiniOpenAI135.065.8—1.05M$0.70
119deepseek-r1-0528-qwen3-8b—9.3———
120GPT-4.1OpenAI136.866.9—1.05M$3.50
121DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
122GPT-4.5 Preview (Feb 2025)OpenAI—68.7———
123Mistral Medium 3Mistral AI134.159.5—131K$0.80
124o1-previewOpenAI134.850.3———
125Gemini 2.0 Flash (Feb 2025)Google134.764.1———
126Mistral Small 3.2Mistral AI131.749.1———
127Magistral Small 1.0Mistral AI133.256.1———
128GPT-4.1 NanoOpenAI129.648.9—1.05M$0.17
129Magistral Small 1.2Mistral AI131.447.6———
130Gemini 1.5 Pro (Sept 2024)Google131.757.2———
131Gemma 3 27BGoogle130.047.7—131K$0.17
132DeepSeek-R1-Distill-Qwen-1.5BDeepSeek—33.6———
133Llama 4 Maverick (FP8)Meta—67.0———
134Qwen Plus (Jan 2025)Alibaba (Qwen)—48.1———
135Gemma 3 12BGoogle123.539.5—131K$0.075
136Gemini 1.5 Flash (Sep 2024)Google129.447.3———
137Qwen2.5-MaxAlibaba (Qwen)132.556.1———
138DeepSeek V3DeepSeek132.456.5—164K$0.45
139Phi 4Microsoft130.456.1—16K$0.087
140Grok-2 (Dec 2024)xAI130.553.8———
141Llama 3.1-405BMeta128.850.9———
142Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
143Claude 3.5 Sonnet (Oct 2024)Anthropic—55.3———
144Qwen2.5-72BAlibaba (Qwen)129.049.1———
145Qwen3-1.7BAlibaba (Qwen)—38.0———
146Llama 4 ScoutMeta129.651.8—1.31M$0.15
147Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
148Gemma 3 4BGoogle116.023.2—131K$0.063
149Qwen2.5-32BAlibaba (Qwen)128.546.1———
150GPT-4o-miniOpenAI126.637.7—128K$0.26
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research