Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 0.0% | 0% | — | 1 Jul 2025 | Epoch ↗ |
| DeepResearch Bench | agents | 46.6% | 84% | 2K | — | External ↗ |
| GSO-Bench | coding | 4.9% | 10% | — | — | External ↗ |
| ARC-AGI-2 | reasoning | 5.9% | 6% | 16K | — | External ↗ |
| METR Time Horizons | agents | 62.0% | 73% | 16K | — | External ↗ |
| GeoBench | multimodal | 37.0% | 42% | — | — | External ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 4.1% ±1.2 | 8% | — | 4 Jul 2025 | Epoch ↗ |
| Fiction.LiveBench | long-context | 46.9% | 48% | — | — | External ↗ |
| Lech Mazur Writing | other | 81.4% | 95% | 16K | — | External ↗ |
| VPCT | multimodal | 34.0% | 37% | 32K | — | External ↗ |
| WeirdML | coding | 46.1% ±0.0 | 49% | 16K | — | External ↗ |
| Aider polyglot | coding | 61.3% | 70% | 32K | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 71.1% ±6.8 | 71% | 32K | 22 May 2025 | Eval log ↗ |
| The Agent Company | agents | 33.1% | 77% | — | — | External ↗ |
| SimpleBench | reasoning | 45.5% | 56% | 12K | — | External ↗ |
| Cybench | coding | 35.0% | 38% | — | — | External ↗ |
| OSWorld | agents | 43.9% | 61% | agent: claude-4-sonnet-20250514 (50 steps) | — | External ↗ |
| GPQA diamond | science | 79.2% ±2.7 | 83% | 59K | 26 May 2025 | Eval log ↗ |
| MATH level 5 | math | 84.4% ±1.0 | 86% | — | 22 May 2025 | Epoch ↗ |
| ARC-AGI | reasoning | 40.0% | 41% | 16K | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
13 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.