science benchmark · included in ECI
Surface Evolver Bench
Surface Evolver Bench benchmark (score column: Mean score).
Results
14
Random baseline
0.0%
Score ceiling
100%
Released
19 Jun 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K3 (unknown) Open source Moonshot | 95.0% | unknown | 16 Jul 2026 | — | Source ↗ | |
| 2 | Kimi K3 Open source MoonshotAI | 93.0% | max | 16 Jul 2026 | — | Source ↗ | |
| 3 | GLM 5.2 Open source Z.ai | 55.6% | high | 16 Jun 2026 | — | Source ↗ | |
| 4 | MiniMax M3 Open source MiniMax | 55.0% | — | 1 Jun 2026 | — | Source ↗ | |
| 5 | GLM 5.3 Flash Open source Z.ai | 52.5% | max | 20 Aug 2026 | — | Source ↗ | |
| 6 | Kimi K2.7 Code Open source MoonshotAI | 48.8% | — | 12 Jun 2026 | — | Source ↗ | |
| 7 | DeepSeek V4.1 Flash Open source DeepSeek | 46.3% | high | 9 Sep 2026 | — | Source ↗ | |
| 8 | Qwen3.8 27B (xhigh) Open source Alibaba | 45.0% | xhigh | 14 Aug 2026 | — | Source ↗ | |
| 9 | Qwen3.6 35B A3B Open source Qwen | 44.4% | — | 27 Apr 2026 | — | Source ↗ | |
| 10 | DeepSeek v4 Pro (high) Open source DeepSeek | 40.0% | high | 24 Apr 2026 | — | Source ↗ | |
| 11 | Gemma 4 31B Open source Google | 30.6% | — | 2 Apr 2026 | — | Source ↗ | |
| 12 | Mistral Medium 3.5 Open source Mistral | 26.9% | — | 28 Apr 2026 | — | Source ↗ | |
| 13 | gpt-oss-120b Open source OpenAI | 25.0% | — | 5 Aug 2025 | — | Source ↗ | |
| 14 | Trinity Large Thinking Open source Arcee AI | 15.6% | — | 1 Apr 2026 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.