Benchmark Radar
AI BENCHMARK PROFILE

OSWorld 2.0

General AIAgentsXLang Lab

Evaluates computer-use agents on 108 long-horizon real-world workflows across everyday and professional tasks, scored by binary completion at 500 steps and partial scores.

Released
2026-06-28
Readiness
Runnable
Primary field
General AI

Why it matters

Captures long-horizon, dynamic, and hidden-state challenges absent in prior benchmarks, revealing that agents fail on constraint tracking and mid-task information.

Motivation

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.