Skip to content
Updated 30m ago

AI Leaderboard - 916 models ranked by capability, price & context

Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.

Highest capability
GPT-6 Astra
ECI 166.6Epoch AI
Leads on reasoning
GPT-6 Astra
95.8% GPQAEpoch AI
Wins at coding
GPT-6 Astra
74.1% DeepSWEEpoch AI
Cheapest in the top 10
GPT-5.6 Sol
$4.00/M blendedOpenRouter
Longest context
Grok 4.20
2.0M tokensOpenRouter
Best open weights
Kimi K3
ECI 157.68Epoch AI

Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
1GPT-6 AstraNEWOpenAI166.695.873.21.05M$20.00
2Claude Fable 5.1NEWAnthropic165.0——1M$20.00
3Claude Fable 5Anthropic163.685.969.71M$20.00
4Claude Opus 5Anthropic162.793.973.61M$10.00
5GPT-5.5 ProOpenAI162.4——1.05M$67.50
6GPT-5.6 SolOpenAI162.093.572.71.05M$4.00
7GPT-5.6 TerraOpenAI159.393.369.61.05M$4.50
8GPT-5.5OpenAI159.3—64.41.05M$11.25
9GPT-5.4 ProOpenAI159.1——1.05M$67.50
10Claude Opus 4.8Anthropic158.391.059.01M$10.00
11Gemini 3.7 FlashGoogle157.794.865.31.05M$1.50
12Kimi K3Moonshot AI157.793.168.51.05M$6.00
13Gemini 3.8 FlashNEWGoogle157.195.473.81.05M$1.50
14GPT-5.4OpenAI156.989.9—1.05M$5.63
15Muse Spark 1.3NEWMeta156.9——1.05M$2.00
16GPT-5.3-CodexOpenAI156.8——400K$4.81
17Qwen 3.8 MaxAlibaba (Qwen)156.7————
18Grok 4.6xAI156.594.065.2500K$3.00
19Claude Opus 4.7Anthropic156.486.4—1M$10.00
20Claude Sonnet 5Anthropic156.380.353.81M$4.00
21GPT-5.6 LunaOpenAI156.391.667.21.05M$0.45
22GLM 5.3Z.ai (Zhipu AI)155.690.969.01.31M$2.15
23GPT-5.2 ProOpenAI155.4——400K$57.75
24DeepSeek V4 Pro 0813DeepSeek155.491.7—1.05M$1.41
25Claude Opus 4.6Anthropic155.488.4—1M$10.00
26Qwen3.8 Max (0902)NEWAlibaba (Qwen)155.3——1M$3.00
27Muse Spark 1.2Meta155.2——1.05M$2.00
28DeepSeek V4.1 FlashNEWDeepSeek155.0——1.05M$0.53
29Gemini 3.1 ProGoogle154.9————
30Gemini 3.5 FlashGoogle154.692.837.41.05M$3.38
31DeepSeek V4 Flash 0731DeepSeek154.591.0—1.31M$0.093
32Gemini 3.6 FlashGoogle154.494.146.71.05M$1.50
33Muse Spark 1.1Meta154.3—53.31.05M$2.00
34Grok 4.5xAI154.093.453.8500K$3.00
35Qwen3.7 MaxAlibaba (Qwen)153.790.9—1M$2.21
36GPT-5.2OpenAI153.488.2—400K$4.81
37Gemini 3 ProGoogle153.0————
38Claude Sonnet 4.6Anthropic152.383.329.91M$6.00
39Muse SparkMeta152.189.8———
40Grok 4.20xAI152.089.3—2M$1.56
41GLM 5.3 FlashZ.ai (Zhipu AI)151.990.263.41.31M$0.24
42Gemini 3 FlashGoogle151.8————
43GLM 5.2Z.ai (Zhipu AI)151.891.943.81.05M$1.39
44Kimi K2.6Moonshot AI151.090.8—262K$1.34
45GPT-5 ProOpenAI150.3——400K$41.25
46Inkling SmallThinking Machines Lab150.1——524K$0.64
47Claude Opus 4.5Anthropic150.1——200K$10.00
48Kimi K2.7 CodeMoonshot AI150.087.930.5262K$1.32
49GPT-5OpenAI150.086.2—400K$3.44
50GLM 5.1Z.ai (Zhipu AI)149.989.9—205K$2.15
— means no published score from that source yet · click a model for every benchmark with its source

New - awaiting independent evaluation

All new models →

Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.

Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research