HopRefusalBench
HopRefusalBench evaluates refusal behavior of search-augmented language model agents on multi-hop questions that are unanswerable, covering three causes of unanswerability and three chain topologies.
- Released
- 2026-08-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Abstention benchmarks typically focus on single-hop queries, leaving evaluation gaps for failures that emerge during multi-hop reasoning and retrieval. This benchmark targets that gap and provides metrics for diagnosing refusal performance, aiding in improving agent reliability.
Motivation
Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question is fundamentally unanswerable remains poorly understood.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.