H2R-Bench
H2R-Bench evaluates cross-embodiment human-to-robot manipulation video generation. Models convert egocentric human demonstrations into robot manipulation videos under specified target embodiments. Scoring covers five dimensions: goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality, aggregated into H2RCore.
- Released
- 2026-08-13
- Readiness
- Runnable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Assesses whether video world models can bridge the embodiment gap between human hands and robotic end-effectors, providing a diagnostic for how well generated videos transfer functional interactions and task execution. Results show generic video quality does not correlate with transfer validity, aiding model selection for robot learning from human video.
Motivation
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.