AI BENCHMARK PROFILE
CAP
CAP evaluates cross-site browser agents on 420 tasks across 108 real websites and 24 domains. Tasks require complex UI interactions and visual perception, with scoring via an agent-as-a-judge framework.
- Released
- 2026-08-09
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
CAP addresses the gap in browser agent evaluation by focusing on cross-site workflows and perception-heavy interactions, which are common in real-world browsing. It provides fine-grained diagnostics to identify bottlenecks in current agents, aiding targeted improvements.
Motivation
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.