Skip to content
agents benchmark · included in ECI

OSWorld

Computer-use tasks in real desktop operating systems.
Results
9
Random baseline
0.0%
Score ceiling
100%
Released
11 Apr 2024

Score by model release date

Best score by organisation

#ModelScoreRelativeEvidence
1Claude Sonnet 4.6 (no thinking)
Anthropic
72.1%Source ↗
2Claude Opus 4.5 (no thinking)
Anthropic
66.3%Source ↗
3Kimi K2.5 Open source
MoonshotAI
63.3%Source ↗
4Claude Sonnet 4.5 (no thinking)
Anthropic
62.9%Source ↗
5Claude Sonnet 4
Anthropic
43.9%Source ↗
6Claude 3.7 Sonnet
Anthropic
35.8%Source ↗
7computer-use-preview-2025-03-11
31.3%Source ↗
8o3 (medium)
OpenAI
23.0%Source ↗
9Qwen2.5-72B Open source
Alibaba
5.0%Source ↗

Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.