HRIBench
HRIBench evaluates intent-aware human-robot collaboration through structured scenario scripts covering roles of Instructor, Collaborator, and Intruder, with 13 tasks and 650 episodes, scoring synchronization, responsiveness, protocol compliance, and safety.
- Released
- 2026-07-05
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing VLA benchmarks focus on isolated manipulation, leaving a gap in evaluating coordination under shared agency. HRIBench provides a standardized protocol for assessing temporal coordination and intent understanding, with evidence that fine-tuning on its data improves real-world task success.
Motivation
Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.