Gemini 3.1 Pro Preview
ModelActiveby Google · family “gemini-pro-preview”
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... (description from the OpenRouter listing)
in: audioin: filein: imagein: textin: videoout: textReasoningTool useStructured output
Epoch Capabilities Index
Not scored by Epoch AI
Released
19 Feb 2026
Context window
1.05M tokens
Max output
66K tokens
Input $ / 1M tokens
$2.00
Output $ / 1M tokens
$12.0
Cached input $ / 1M
$0.20
Each benchmark shown separately with its own source
Benchmark results (32)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Furniture Assembly | multimodal | 26.7% ±5.6 | 32% | high | 10 Sep 2026 | Epoch ↗ |
| LMCA | agents | 53.8% | 79% | high | — | External ↗ |
| DTBench | reasoning | 97.1% | 98% | medium | — | External ↗ |
| Mystery Game Puzzles | games | 34.0% ±4.8 | 40% | high | 27 Jul 2026 | Epoch ↗ |
| EBR-bench | reasoning | 14.3% | 19% | — | 25 Jun 2026 | Epoch ↗ |
| MirrorCode | coding | 8.9% ±4.6 | 12% | high | 10 Aug 2026 | Epoch ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 59.6% ±2.9 | 64% | — | 11 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 26.8% ±7.0 | 27% | — | 11 Jun 2026 | Eval log ↗ |
| DeepSWE | coding | 11.7% | 16% | — | — | External ↗ |
| ExploitBench | coding | 26.1% | 35% | — | — | External ↗ |
| CL-bench Life | long-context | 16.9% | 76% | — | — | External ↗ |
| PostTrainBench | agents | 22.0% | 53% | — | — | External ↗ |
| CL-bench | long-context | 20.8% | 75% | — | — | External ↗ |
| ProofBench | math | 26.0% | 26% | — | — | External ↗ |
| APEX-Agents | agents | 35.3% | 47% | — | — | External ↗ |
| Chess Puzzles | games | 55.0% ±5.0 | 76% | — | 19 Feb 2026 | Eval log ↗ |
| SimpleQA Verified | knowledge | 73.5% ±1.4 | 97% | high | 10 Aug 2026 | Epoch ↗ |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 16.7% ±5.4 | 35% | — | 19 Feb 2026 | Epoch ↗ |
| DeepResearch Bench | agents | 47.8% | 86% | high | — | External ↗ |
| GSO-Bench | coding | 22.6% | 48% | — | — | External ↗ |
| Terminal Bench | agents | 80.2% ±2.6 | 95% | agent: TongAgents | 19 Feb 2026 | External ↗ |
| ARC-AGI-2 | reasoning | 77.1% | 81% | — | — | External ↗ |
| METR Time Horizons | agents | 77.0% | 90% | — | — | External ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 36.9% ±2.8 | 70% | — | 19 Feb 2026 | Epoch ↗ |
| HLE | knowledge | 46.4% | 85% | — | — | External ↗ |
| WeirdML | coding | 72.1% ±0.0 | 77% | — | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 95.6% ±3.1 | 96% | — | 20 Feb 2026 | Eval log ↗ |
| Balrog | games | 57.0% | 83% | — | — | External ↗ |
| SimpleBench | reasoning | 79.6% | 97% | — | — | External ↗ |
| SWE-Bench verified | coding | 75.6% ±2.0 | 91% | — | 24 Feb 2026 | Eval log ↗ |
| GPQA diamond | science | 94.4% ±1.6 | 99% | high | 6 Aug 2026 | Eval log ↗ |
| ARC-AGI | reasoning | 98.0% | 99% | — | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
API list price over time
Price history
$ per 1M tokens
8 recorded prices on 3 Mar 2026 (OpenRouter listing + Internet Archive snapshots); re-read on every data refresh, most recently 7h ago. Steps show when the price changed.
Events
- BenchmarkEpoch AI evaluates Gemini 3.1 Pro PreviewEpoch AI Benchmarking Hub
- Model launchMajorGoogle releases Gemini 3.1 Pro PreviewOpenRouter models API