Skip to content

Reasoning models

Reasoning models spend extra computation "thinking" through intermediate steps before answering, which improves results on maths, coding and multi-step problems at the cost of speed and tokens.

Many providers offer the same model at several reasoning-effort settings (for example low, medium, high). Higher settings are usually more accurate on hard problems but slower and more expensive, because the thinking tokens are billed too.

On this site, those settings are folded into one row per model on the leaderboards, showing the best-scoring setting.

Related ranking: Best for reasoning

  1. 1.GPT-6 Astra OpenAI95.8%
  2. 2.Claude Sonnet 5.5 Anthropic95.6%
  3. 3.Gemini 3.8 Flash Google95.4%
  4. 4.Gemini 3.7 Flash Google94.8%
  5. 5.GPT-5.4 Pro OpenAI94.6%

Ranked by GPQA Diamond - graduate-level science questions written to be hard to look up. Source: Epoch AI. Full ranking

Related terms

Explanation written by AI Stats Live editors; last reviewed 29 Sep 2026. Examples and rankings update automatically from the sources listed on the methodology page.