GPT-5.4
ModelActiveby OpenAI · family “gpt”
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for... (description from the OpenRouter listing)
in: textin: imagein: fileout: textReasoningTool useStructured output
Epoch Capabilities Index
156.9 #14 of 268
90% CI 155.1 – 159.3
Released
5 Mar 2026
Context window
1.05M tokens
Max output
128K tokens
Input $ / 1M tokens
$2.50
Output $ / 1M tokens
$15.0
Cached input $ / 1M
$0.25
Each benchmark shown separately with its own source
Benchmark results (10)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| MirrorCode | coding | 15.6% ±7.7 | 20% | high | 10 Aug 2026 | Epoch ↗ |
| CL-bench Life | long-context | 19.3% | 87% | high | — | External ↗ |
| PostTrainBench | agents | 19.0% | 45% | high | — | External ↗ |
| Chess Puzzles | games | 38.0% ±4.9 | 53% | high | 15 Jul 2026 | Eval log ↗ |
| GSO-Bench | coding | 25.5% | 54% | high | — | External ↗ |
| ARC-AGI-2 | reasoning | 67.5% | 71% | high | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 97.8% ±2.2 | 98% | high | 15 Jul 2026 | Eval log ↗ |
| SWE-Bench verified | coding | 76.9% ±1.9 | 92% | high | 6 Mar 2026 | Eval log ↗ |
| GPQA diamond | science | 89.9% ±2.1 | 94% | high | 15 Jul 2026 | Eval log ↗ |
| ARC-AGI | reasoning | 92.7% | 94% | high | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
API list price over time
Price history
$ per 1M tokens
7 recorded prices on 1 Apr 2026 (OpenRouter listing + Internet Archive snapshots); re-read on every data refresh, most recently 40m ago. Steps show when the price changed.
Events
- Model launchMajorOpenAI releases GPT-5.4Epoch AI Benchmarking Hub