Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Furniture Assembly | multimodal | 28.3% ±5.9 | 34% | max | 10 Sep 2026 | Epoch ↗ |
| LMCA | agents | 55.8% | 85% | max | — | External ↗ |
| DTBench | reasoning | 91.2% | 93% | max | — | External ↗ |
| Mystery Game Puzzles | games | 25.0% ±4.4 | 30% | max | 25 Jul 2026 | Epoch ↗ |
| EBR-bench | reasoning | 12.7% | 17% | max | 29 Jun 2026 | Epoch ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 66.0% ±2.8 | 70% | max | 11 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 26.8% ±7.0 | 27% | max | 11 Jun 2026 | Eval log ↗ |
| CL-bench Life | long-context | 17.0% | 77% | high | — | External ↗ |
| ProofBench | math | 50.0% | 50% | max | — | External ↗ |
| APEX-Agents | agents | 46.3% | 67% | max | — | External ↗ |
| Chess Puzzles | games | 14.0% ±3.5 | 19% | max | 6 Aug 2026 | Eval log ↗ |
| SimpleQA Verified | knowledge | 47.0% ±1.6 | 62% | max | 27 Aug 2026 | Eval log ↗ |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 22.9% ±6.1 | 48% | max | 12 Feb 2026 | Epoch ↗ |
| DeepResearch Bench | agents | 55.3% | 100% | high | — | External ↗ |
| GSO-Bench | coding | 41.2% | 88% | high | — | External ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 40.7% ±2.9 | 78% | max | 12 Feb 2026 | Epoch ↗ |
| HLE | knowledge | 34.4% | 63% | max | — | External ↗ |
| WeirdML | coding | 78.0% ±0.0 | 83% | high | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 91.1% ±4.3 | 91% | max | 6 Aug 2026 | Eval log ↗ |
| GPQA diamond | science | 88.4% ±2.3 | 92% | max | 6 Aug 2026 | Eval log ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
8 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.