agents benchmark · included in ECI
OSWorld
Computer-use tasks in real desktop operating systems.
Results
2
Random baseline
0.0%
Score ceiling
100%
Released
11 Apr 2024
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K2.5 Open source MoonshotAI | 63.3% | agent: Kimi-K2.5 | 27 Jan 2026 | — | OS World Website ↗ | |
| 2 | Qwen2.5-72B Open source Alibaba | 5.0% | agent: qwen2.5-vl-72b-instruct (100 steps) | 19 Sep 2024 | — | OS World Website ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.