WatchAct
WatchAct evaluates robot manipulation from observed human behavior. Each instance pairs a human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task. It covers 3,000 long-horizon instances across 14 tasks in four capability domains: Event Grounding, Procedural Reasoning, Implicit Intent Inference, and Episodic Reasoning. The evaluation protocol separately measures video-to-plan reasoning, policy execution under oracle plans, and full task completion, in simulation and on a Franka Research 3 robot.
- Released
- 2026-06-24
- Readiness
- Inspectable
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing manipulation benchmarks typically evaluate from a single current image, lacking grounding in observed human behavior. WatchAct fills this gap by assessing robots' ability to reason about events, procedures, intents, and scene changes from video, which is critical for real-world human-robot collaboration. It provides a disentangled evaluation to isolate reasoning and execution failures, offering practical value for developing and comparing systems.
Motivation
A robot working alongside people must reason about what they have done, in what order, and with what intent.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.