Updated 46m ago
AI Leaderboard - 916 models ranked by capability, price & context
Independent benchmark results from Epoch AI and Artificial Analysis, live API prices and context windows from OpenRouter. Each ranking uses one named measure - we don't blend them into a made-up score.
Ranked by the Epoch Capabilities Index (ECI) - Epoch AI's statistical model over dozens of benchmarks. Source: Epoch AI. Variants (reasoning effort, batch) are folded into one row.
| # | Model | Capability | Reasoning | Coding | Context | $/M | Weights | Released |
|---|---|---|---|---|---|---|---|---|
| 251 | CodeQwen1.5-7B | 94.2 | — | — | — | — | Closed | 15 Apr 2024 |
| 252 | vicuna-13b-v1.1 | 94.0 | — | — | — | — | Closed | 12 Apr 2023 |
| 253 | MPT-7BMosaicML | 94.0 | — | — | — | — | Open | 5 May 2023 |
| 254 | Gemma 2BGoogle | 93.6 | — | — | — | — | Open | 21 Feb 2024 |
| 255 | StarCoder 2 7BHugging Face | 92.9 | — | — | — | — | Open | 20 Feb 2024 |
| 256 | XGen-7BSalesforce | 92.8 | — | — | — | — | Open | 27 Jun 2023 |
| 257 | Qwen-1_8B | 92.3 | — | — | — | — | Closed | 30 Nov 2023 |
| 258 | open_llama_7b | 91.0 | — | — | — | — | Closed | 7 Jun 2023 |
| 259 | Phi-1.5Microsoft | 90.8 | — | — | — | — | Open | 11 Sep 2023 |
| 260 | Baichuan1-7BBaichuan | 89.8 | — | — | — | — | Open | 1 Jun 2023 |
| 261 | RedPajama-INCITE-7B-Base | 89.7 | — | — | — | — | Closed | 4 May 2023 |
| 262 | Dolly 2.0-12bDatabricks | 88.9 | — | — | — | — | Open | 11 Apr 2023 |
| 263 | DeepSeek Coder 6.7BDeepSeek | 88.9 | — | — | — | — | Open | 2 Nov 2023 |
| 264 | StarCoder 2 3BHugging Face | 88.0 | — | — | — | — | Open | 22 Feb 2024 |
| 265 | Qwen2.5-Coder-0.5B | 87.4 | — | — | — | — | Closed | 18 Sep 2024 |
| 266 | Cerebras-GPT-13BCerebras | 82.4 | — | — | — | — | Open | 20 Mar 2023 |
| 267 | DeepSeek Coder 1.3BDeepSeek | 62.2 | — | — | — | — | Open | 2 Nov 2023 |
| 268 | stablelm-tuned-alpha-7b | 54.3 | — | — | — | — | Closed | 19 Apr 2023 |
— means no published score from that source yet · click a model for every benchmark with its source
New - awaiting independent evaluation
All new models →- Claude Sonnet 5.5Anthropic · 28 Sep 2026 · AA 56 · 1M · $4.00/M
- Qwen3.8 Max PrimeAlibaba (Qwen) · 23 Sep 2026 · 1M · $6.00/M
- GPT-6 SolOpenAI · 22 Sep 2026 · AA 47.5 · 1.05M · $4.00/M
- GPT-6 LunaOpenAI · 22 Sep 2026 · AA 37.3 · 1.05M · $0.20/M
- GPT-6 Sol ProOpenAI · 22 Sep 2026 · 1.05M · $4.00/M
- GPT-6 Luna ProOpenAI · 22 Sep 2026 · 1.05M · $0.20/M
- Claude Opus 5.5Anthropic · 22 Sep 2026 · AA 57.6 · 1M · $8.00/M
- Grok 4.7xAI · 21 Sep 2026 · AA 46.4 · 500K · $3.00/M
Not ranked until an independent evaluator (Epoch AI) publishes scores. We don't use vendor-reported benchmark claims for rankings.
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research