MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Surface Evolver Bench | science | 48.8% | 51% | — | — | External ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 54.0% ±3.0 | 58% | — | 13 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 12.2% ±5.2 | 12% | — | 13 Jun 2026 | Eval log ↗ |
| FrontierCode | coding | 30.1% | 56% | — | — | External ↗ |
| DeepSWE | coding | 30.5% | 41% | — | — | External ↗ |
| APEX-Agents | agents | 37.6% | 55% | — | — | External ↗ |
| Chess Puzzles | games | 21.0% ±4.1 | 29% | — | 7 Aug 2026 | Eval log ↗ |
| SimpleQA Verified | knowledge | 36.5% ±1.5 | 48% | — | 27 Aug 2026 | Eval log ↗ |
| WeirdML | coding | 54.1% ±0.0 | 58% | — | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 95.6% ±3.1 | 96% | — | 7 Aug 2026 | Eval log ↗ |
| SimpleBench | reasoning | 57.9% | 71% | — | — | External ↗ |
| GPQA diamond | science | 87.9% ±2.3 | 92% | — | 7 Aug 2026 | Eval log ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
4 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.