Benchmark Radar
AI BENCHMARK PROFILE

WeaveBench

General AIKnowledge & ReasoningWeaveBench team

WeaveBench evaluates computer-use agents on 114 long-horizon tasks across 8 real-world work domains. Each task interleaves GUI interaction with command-line and code operations in a single trajectory, with scoring based on a trajectory-aware judge that detects fabricated evidence.

Released
2026-06-08
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks often evaluate interfaces separately, leaving hybrid orchestration under-tested. WeaveBench fills this gap by requiring agents to combine GUI and CLI/code in realistic tasks, and exposes that outcome-only grading overestimates performance.

Motivation

Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.