WeaveBench
WeaveBench evaluates computer-use agents on 114 long-horizon tasks across 8 real-world work domains. Each task interleaves GUI interaction with command-line and code operations in a single trajectory, with scoring based on a trajectory-aware judge that detects fabricated evidence.
- Released
- 2026-06-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks often evaluate interfaces separately, leaving hybrid orchestration under-tested. WeaveBench fills this gap by requiring agents to combine GUI and CLI/code in realistic tasks, and exposes that outcome-only grading overestimates performance.
Motivation
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.