Benchmark Radar
AI BENCHMARK PROFILE

HeraBench

Consumer & ProductivityAgents

HeraBench is a fault-injected benchmark for multi-device agent workflows on Linux and Android, evaluating hierarchical replanning under injected strategy- and device-level failures.

Released
2026-06-18
Readiness
Paper only
Primary field
Consumer & Productivity

Why it matters

Current multi-device agent benchmarks lack systematic fault injection to test recovery capabilities, making it difficult to compare hierarchical versus global replanning approaches.

Motivation

Real-world computer-use tasks often span multiple applications and devices, requiring agents to coordinate heterogeneous environments under dynamic runtime failures.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.