Skip to content
Updated 45m ago

AI Leaderboard - 916 models ranked by capability, price & context

Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.

Highest capability
GPT-6 Astra
ECI 166.6Epoch AI
Leads on reasoning
GPT-6 Astra
95.8% GPQAEpoch AI
Wins at coding
GPT-6 Astra
74.1% DeepSWEEpoch AI
Cheapest in the top 10
GPT-5.6 Sol
$4.00/M blendedOpenRouter
Longest context
Grok 4.20
2.0M tokensOpenRouter
Best open weights
Kimi K3
ECI 157.68Epoch AI

Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
101Gemma 4 26B A4BGoogle141.873.2—262K$0.14
102Claude Sonnet 4Anthropic141.779.2—200K$6.00
103Gemini 2.5 Flash (May 2025)Google141.5————
104Mistral Medium 3.5Mistral AI141.4——262K$3.00
105DeepSeek-R1 (May 2025)DeepSeek141.376.3———
106Claude 3.7 SonnetAnthropic141.266.0———
107Gemini 2.5 Flash (Jun 2025)Google140.8————
108Grok-3 minixAI140.376.3———
109o3 MiniOpenAI140.377.0—200K$1.93
110Kimi K2 (Jul 2025)Moonshot AI140.1————
111Gemini 2.5 Flash (Apr 2025)Google140.0————
112gpt-oss-120bOpenAI139.975.8—131K$0.070
113DeepSeek V3.1DeepSeek139.9——164K$0.42
114Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba (Qwen)139.670.1———
115Qwen3.5-9BAlibaba (Qwen)139.479.0—262K$0.11
116GPT-5 NanoOpenAI139.469.4—400K$0.14
117Qwen3 235B A22BAlibaba (Qwen)139.470.7—131K$0.80
118R1DeepSeek139.071.7—64K$1.15
119Qwen3-235B-A22B-Instruct (Jul 2025)Alibaba (Qwen)138.9————
120Qwen3 32BAlibaba (Qwen)138.565.7—131K$0.13
121Grok 3xAI138.375.8———
122Qwen3 14BAlibaba (Qwen)138.263.8—131K$0.15
123gpt-oss-20bOpenAI137.860.8—131K$0.036
124QwQ-32BAlibaba (Qwen)137.665.3———
125DeepSeek-R1-Distill-Qwen-32BDeepSeek137.464.1———
126Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba (Qwen)137.455.6———
127GPT-4.1OpenAI136.866.9—1.05M$3.50
128GPT-4.5OpenAI136.7————
129Qwen3 30B A3BAlibaba (Qwen)136.261.7—131K$0.23
130Qwen3 8BAlibaba (Qwen)136.256.8—131K$0.20
131DeepSeek-V3 (Mar 2025)DeepSeek135.967.6———
132o1-miniOpenAI135.862.4———
133DeepSeek-R1-Distill-Qwen-14BDeepSeek135.444.7———
134Gemini 2.0 Flash Thinking (Jan 2025)Google135.4————
135Gemini 2.0 ProGoogle135.1————
136GPT-4.1 MiniOpenAI135.065.8—1.05M$0.70
137o1-previewOpenAI134.850.3———
138Gemini 2.0 Flash (Dec 2024)Google134.7————
139Gemini 2.0 Flash (Feb 2025)Google134.764.1———
140Mistral Medium 3Mistral AI134.159.5—131K$0.80
141Gemini 2.5 Flash-Lite (Jun 2025)Google133.9————
142Claude 3.5 Sonnet (October 2024)Anthropic133.6————
143Magistral Small 1.0Mistral AI133.256.1———
144Qwen2.5-MaxAlibaba (Qwen)132.556.1———
145DeepSeek V3DeepSeek132.456.5—164K$0.45
146Llama 4 MaverickMeta132.2——1.05M$0.30
147Gemini 1.5 Pro (Sept 2024)Google131.757.2———
148Mistral Small 3.2Mistral AI131.749.1———
149Magistral Small 1.2Mistral AI131.447.6———
150Grok-2 (Dec 2024)xAI130.553.8———
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research