DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| ARC-AGI-2 | reasoning | 1.3% | 1% | — | — | External ↗ |
| METR Time Horizons | agents | 51.9% | 61% | — | — | External ↗ |
| Fiction.LiveBench | long-context | 69.4% | 71% | — | — | External ↗ |
| Lech Mazur Writing | other | 83.0% | 97% | — | — | External ↗ |
| WeirdML | coding | 36.5% ±0.0 | 39% | — | — | External ↗ |
| Aider polyglot | coding | 56.9% | 65% | — | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 53.3% ±7.5 | 53% | — | 26 Feb 2025 | Eval log ↗ |
| Balrog | games | 34.9% | 51% | — | — | External ↗ |
| SimpleBench | reasoning | 30.9% | 38% | — | — | External ↗ |
| GPQA diamond | science | 71.7% ±3.1 | 75% | — | 26 May 2025 | Eval log ↗ |
| MATH level 5 | math | 93.1% ±0.7 | 95% | — | 31 Jan 2025 | Epoch ↗ |
| ARC-AGI | reasoning | 15.8% | 16% | — | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
13 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.