AI BENCHMARK PROFILE
WebRetriever
Introduces a benchmark with 800 websites and 1,550 tasks for web agent evaluation, plus the NavEval LLM-as-Judge framework and three evaluation protocols.
- Released
- 2026-07-07
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Could offer large-scale cross-domain assessment for web agents, but the evaluation methodology relies on LLM-as-Judge and lacks clear public implementation details.
Motivation
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.