Skip to content

Best for reasoning

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up.

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
201Ministral 3BMistral AI118.125.3———
202DeepSeek LLM 67BDeepSeek110.524.6———
203granite-4.0-1b—24.0———
204Llama 3.2 1BMeta102.023.9———
205Gemma 3 4BGoogle116.023.2—131K$0.063
206Gemma 3 1BGoogle—19.9———
207Mistral 7B v0.3Mistral AI108.715.2———
208Yi-34B01.AI117.314.7———
209granite-4.0-350m—11.2———
210deepseek-r1-0528-qwen3-8b—9.3———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research