Skip to content

Best for reasoning

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up.

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba (Qwen)139.670.1———
102GPT-5 NanoOpenAI · +3 variants139.469.4—400K$0.14
103GPT-4.5 Preview (Feb 2025)OpenAI—68.7———
104DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
105Llama 4 Maverick (FP8)Meta—67.0———
106GPT-4.1OpenAI136.866.9—1.05M$3.50
107GPT-4.1 MiniOpenAI135.065.8—1.05M$0.70
108Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
109Gemini 2.0 Pro Exp (Feb 2025)Google—65.7———
110QWQ-PlusAlibaba (Qwen)—65.4———
111QwQ-32BAlibaba (Qwen)137.665.3———
112DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
113Gemini 2.0 Flash (Feb 2025)Google134.764.1———
114Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
115o1-miniOpenAI · +1 variant135.862.4———
116Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
117gpt-oss-20bOpenAI137.860.8—131K$0.036
118GLM 4.7 FlashZ.ai (Zhipu AI)—60.5—200K$0.15
119Mistral Medium 3Mistral AI134.159.5—131K$0.80
120Gemini 1.5 Pro (Sept 2024)Google131.757.2———
121Gemini 2.0 Flash Thinking ExpGoogle—57.1———
122Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
123DeepSeek V3DeepSeek132.456.5—164K$0.45
124Qwen2.5-MaxAlibaba (Qwen)132.556.1———
125Magistral Small 1.0Mistral AI133.256.1———
126Phi 4Microsoft130.456.1—16K$0.087
127DeepSeek-R1-Distill-Llama-70BDeepSeek—55.7———
128Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
129Claude 3.5 Sonnet (Oct 2024)Anthropic—55.3———
130Claude 3.5 Sonnet (Jun 2024)Anthropic—54.0———
131Grok-2 (Dec 2024)xAI130.553.8———
132Qwen3-4BAlibaba (Qwen)—52.3———
133Llama 4 ScoutMeta129.651.8—1.31M$0.15
134Mistral Large 2 (Nov 2024)Mistral AI128.551.3———
135Llama 3.1-405BMeta128.850.9———
136o1-previewOpenAI134.850.3———
137GPT-4o (Aug 2024)OpenAI128.849.2———
138Qwen2.5-72BAlibaba (Qwen)129.049.1———
139Mistral Small 3.2Mistral AI131.749.1———
140Mistral Large 2 (Jul 2024)Mistral AI127.549.0———
141GPT-4.1 NanoOpenAI129.648.9—1.05M$0.17
142GPT-4o (May 2024)OpenAI129.048.9———
143Qwen Plus (Jan 2025)Alibaba (Qwen)—48.1———
144GPT-4o (Nov 2024)OpenAI128.847.9———
145Gemma 3 27BGoogle130.047.7—131K$0.17
146Magistral Small 1.2Mistral AI131.447.6———
147Mistral Small 3.1Mistral AI127.547.5———
148Llama 3.3 70BMeta127.347.4———
149Gemini 1.5 Flash (Sep 2024)Google129.447.3———
150Mistral Small 3Mistral AI127.147.3—33K$0.058
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research