Benchmark Radar
AI BENCHMARK PROFILE

EvoBrowseComp

General AIKnowledge & Reasoning

EvoBrowseComp evaluates search agents on evolving knowledge with 800 contamination-free questions synthesized from live web traversal. The benchmark is designed to be regularly updated to prevent contamination.

Released
2026-06-11
Readiness
Paper only
Primary field
General AI

Why it matters

Static benchmarks suffer from contamination and memorization, obscuring genuine retrieval. EvoBrowseComp offers a dynamic, auto-updated evaluation, which is critical for measuring true browsing competence.

Motivation

Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.