Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| LMCA | agents | 46.5% | 71% | max | — | External ↗ |
| DTBench | reasoning | 89.9% | 91% | max | — | External ↗ |
| OSWorld 2.0 | agents | 8.3% | 26% | max | — | External ↗ |
| DeepSWE | coding | 29.9% | 40% | high | — | External ↗ |
| ProofBench | math | 45.0% | 45% | max | — | External ↗ |
| APEX-Agents | agents | 43.0% | 63% | high | — | External ↗ |
| Chess Puzzles | games | 5.0% ±2.2 | 7% | high | 13 Jul 2026 | Eval log ↗ |
| SimpleQA Verified | knowledge | 35.5% ±1.5 | 47% | high | 10 Aug 2026 | Epoch ↗ |
| DeepResearch Bench | agents | 54.9% | 99% | high | — | External ↗ |
| ARC-AGI-2 | reasoning | 60.4% | 64% | high | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 75.6% ±6.5 | 76% | high | 13 Jul 2026 | Eval log ↗ |
| GPQA diamond | science | 83.3% ±2.7 | 87% | high | 13 Jul 2026 | Eval log ↗ |
| ARC-AGI | reasoning | 86.5% | 88% | high | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
8 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.