AI BENCHMARK PROFILE
OSWorld 2.0
Evaluates computer-use agents on 108 long-horizon real-world workflows across everyday and professional tasks, scored by binary completion at 500 steps and partial scores.
- Released
- 2026-06-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Captures long-horizon, dynamic, and hidden-state challenges absent in prior benchmarks, revealing that agents fail on constraint tracking and mid-task information.
Motivation
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.