AI BENCHMARK PROFILE
EvoBrowseComp
EvoBrowseComp evaluates search agents on evolving knowledge with 800 contamination-free questions synthesized from live web traversal. The benchmark is designed to be regularly updated to prevent contamination.
- Released
- 2026-06-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Static benchmarks suffer from contamination and memorization, obscuring genuine retrieval. EvoBrowseComp offers a dynamic, auto-updated evaluation, which is critical for measuring true browsing competence.
Motivation
Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.