Skip to content

Newest models

Most recent releases, with whatever independent scores exist so far.

Most recent releases, with whatever independent scores exist so far. Source: OpenRouter / Epoch AI. Variants (reasoning effort, batch) are folded into one row.

#ModelCapabilityReasoningCodingContext$/M
651Dolly 2.0-12bDatabricks88.9————
652Cerebras-GPT-13BCerebras82.4————
653GPT-4 (Mar 2023)OpenAI125.935.7———
654LLaMA-65BMeta109.9————
655LLaMA-33BMeta107.1————
656LLaMA-13BMeta100.1————
657LLaMA-7BMeta96.1————
658BLIP-2 (Q-Former)Salesforce Research—————
— means no published score from that source yet · click a model for every benchmark with its source
Best AI for coding →
Models ranked on DeepSWE
Best AI for reasoning →
Ranked on GPQA Diamond
Best AI for math →
Ranked on competition math
Best open-weight models →
Weights you can run yourself
Cheapest capable models →
Lowest blended API price
Best AI coding tools →
Apps & agents by popularity
Best AI image generators →
Apps & models
Best AI for research →
Answer engines & deep research