Benchmark Radar
AI BENCHMARK PROFILE

IHBench

General AIMultimodal Perception

IHBench evaluates post-interruption recovery in voice agents executing state-machine-driven workflows across 10 enterprise domains. It scores task fulfillment and recovery quality for six interruption types.

Released
2026-06-17
Readiness
Paper only
Primary field
General AI

Why it matters

Voice agents must handle interruptions while maintaining workflow progress, but existing benchmarks measure only interruption timing. IHBench focuses on recovery quality, a distinct capability axis important for deployed agents.

Motivation

Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress through multi-step procedures.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.