Claude Opus 4.8
ModelActiveby Anthropic · family “claude-opus”
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token... (description from the OpenRouter listing)
in: textin: imagein: fileout: textReasoningTool useStructured output
Epoch Capabilities Index
158.3 #10 of 268
90% CI 156.3 – 161.0
Released
28 May 2026
Context window
1M tokens
Max output
128K tokens
Input $ / 1M tokens
$5.00
Output $ / 1M tokens
$25.0
Cached input $ / 1M
$0.50
Each benchmark shown separately with its own source
Benchmark results (27)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Furniture Assembly | multimodal | 42.5% ±6.4 | 51% | max | 10 Sep 2026 | Epoch ↗ |
| LMCA | agents | 57.5% | 84% | max | — | External ↗ |
| DTBench | reasoning | 94.9% | 96% | max | — | External ↗ |
| Mystery Game Puzzles | games | 36.0% ±4.8 | 43% | max | 25 Jul 2026 | Epoch ↗ |
| EBR-bench | reasoning | 28.6% ±1.4 | 38% | max | 7 Aug 2026 | Epoch ↗ |
| OSWorld 2.0 | agents | 20.6% | 66% | max | — | External ↗ |
| Surface Evolver Bench | science | 87.5% | 92% | high | — | External ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 80.0% ±2.4 | 85% | max | 10 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 56.1% ±7.8 | 57% | max | 10 Jun 2026 | Eval log ↗ |
| FrontierCode | coding | 46.5% | 87% | unknown | — | External ↗ |
| DeepSWE | coding | 59.0% | 80% | max | — | External ↗ |
| PostTrainBench | agents | 33.8% | 81% | high | — | External ↗ |
| ProofBench | math | 69.0% | 69% | max | — | External ↗ |
| APEX-Agents | agents | 48.9% | 65% | max | — | External ↗ |
| Chess Puzzles | games | 34.0% ±4.8 | 47% | max | 29 May 2026 | Epoch ↗ |
| Remote Labor Index | agents | 8.3% | 40% | unknown | — | External ↗ |
| SimpleQA Verified | knowledge | 53.0% ±1.6 | 70% | max | 27 Aug 2026 | Eval log ↗ |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 31.3% ±6.8 | 65% | max | 8 Jun 2026 | Epoch ↗ |
| DeepResearch Bench | agents | 50.2% | 91% | high | — | External ↗ |
| GSO-Bench | coding | 47.1% | 100% | unknown | — | External ↗ |
| ARC-AGI-2 | reasoning | 72.1% | 76% | high | — | External ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 47.2% ±2.9 | 90% | max | 8 Jun 2026 | Epoch ↗ |
| WeirdML | coding | 82.9% ±0.0 | 89% | xhigh | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 98.3% ±1.4 | 98% | max | 7 Jun 2026 | Epoch ↗ |
| SimpleBench | reasoning | 64.8% | 79% | unknown | — | External ↗ |
| GPQA diamond | science | 91.0% ±1.9 | 95% | max | 7 Jun 2026 | Epoch ↗ |
| ARC-AGI | reasoning | 92.5% | 94% | max | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
API list price over time
Price history
$ per 1M tokens
5 recorded prices on 6 Jun 2026 (OpenRouter listing + Internet Archive snapshots); re-read on every data refresh, most recently 7h ago. Steps show when the price changed.
Events
- BenchmarkEpoch AI evaluates Claude Opus 4.8Epoch AI Benchmarking Hub
- BenchmarkEpoch AI evaluates Claude Opus 4.8Epoch AI Benchmarking Hub
- Model launchMajorAnthropic releases Claude Opus 4.8Epoch AI Benchmarking Hub