DiscoBench
DiscoBench evaluates search agents on clarification-aware deep search, covering 211 samples and 463 ambiguity instances across 11 domains, with four ambiguity types, measuring task utility, ambiguity detection, interaction strategy, and cost efficiency.
- Released
- 2026-06-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark addresses the gap in evaluating search agents' ability to handle ambiguous and underspecified queries, which is common in real-world search. It assesses proactive clarification and interaction efficiency, offering practical value for improving agent decision-making in complex information-seeking tasks.
Motivation
Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval and reasoning to fulfill user goals.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.