Benchmark Radar
AI BENCHMARK PROFILE

WebRetriever

General AIMultimodal Perception

Introduces a benchmark with 800 websites and 1,550 tasks for web agent evaluation, plus the NavEval LLM-as-Judge framework and three evaluation protocols.

Released
2026-07-07
Readiness
Paper only
Primary field
General AI

Why it matters

Could offer large-scale cross-domain assessment for web agents, but the evaluation methodology relies on LLM-as-Judge and lacks clear public implementation details.

Motivation

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.