AI BENCHMARK PROFILE
RegretBench
RegretBench evaluates clarification policies in multi-turn conversational LLMs, using hidden-intent tasks and a regret-based objective to measure value loss relative to a reference policy.
- Released
- 2026-07-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses the evaluation gap in conversational AI by jointly measuring intent resolution, interaction cost, and stopping decisions, offering a more comprehensive assessment of clarification behavior.
Motivation
Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.