Benchmark Radar
AI BENCHMARK PROFILE

DiscoBench

General AIKnowledge & Reasoning

DiscoBench evaluates search agents on clarification-aware deep search, covering 211 samples and 463 ambiguity instances across 11 domains, with four ambiguity types, measuring task utility, ambiguity detection, interaction strategy, and cost efficiency.

Released
2026-06-26
Readiness
Paper only
Primary field
General AI

Why it matters

This benchmark addresses the gap in evaluating search agents' ability to handle ambiguous and underspecified queries, which is common in real-world search. It assesses proactive clarification and interaction efficiency, offering practical value for improving agent decision-making in complex information-seeking tasks.

Motivation

Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval and reasoning to fulfill user goals.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.