GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| FrontierMath-Tiers-1-3-v2-Private | math | 36.8% ±2.9 | 39% | — | 29 Aug 2026 | Eval log ↗ |
| ExploitBench | coding | 18.1% | 25% | — | — | External ↗ |
| ProofBench | math | 22.2% | 22% | — | — | External ↗ |
| APEX-Agents | agents | 40.9% | 60% | — | — | External ↗ |
| Chess Puzzles | games | 19.0% ±3.9 | 26% | — | 10 Aug 2026 | Eval log ↗ |
| SimpleQA Verified | knowledge | 34.0% ±1.5 | 45% | — | 27 Aug 2026 | Eval log ↗ |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 12.5% ±4.8 | 26% | — | 12 May 2026 | Epoch ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 33.4% ±2.8 | 64% | — | 11 May 2026 | Epoch ↗ |
| WeirdML | coding | 57.1% ±0.0 | 61% | — | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 93.3% ±3.8 | 93% | — | 10 Aug 2026 | Eval log ↗ |
| SimpleBench | reasoning | 55.1% | 67% | — | — | External ↗ |
| SWE-Bench verified | coding | 74.2% ±2.0 | 89% | — | 15 May 2026 | Epoch ↗ |
| GPQA diamond | science | 89.9% ±2.1 | 94% | — | 10 Aug 2026 | Eval log ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
8 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.