Benchmark Radar
AI BENCHMARK PROFILE

HopRefusalBench

General AIKnowledge & Reasoning

HopRefusalBench evaluates refusal behavior of search-augmented language model agents on multi-hop questions that are unanswerable, covering three causes of unanswerability and three chain topologies.

Released
2026-08-02
Readiness
Paper only
Primary field
General AI

Why it matters

Abstention benchmarks typically focus on single-hop queries, leaving evaluation gaps for failures that emerge during multi-hop reasoning and retrieval. This benchmark targets that gap and provides metrics for diagnosing refusal performance, aiding in improving agent reliability.

Motivation

Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question is fundamentally unanswerable remains poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.