Benchmark Radar
AI BENCHMARK PROFILE

WatchAct

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

WatchAct evaluates robot manipulation from observed human behavior. Each instance pairs a human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task. It covers 3,000 long-horizon instances across 14 tasks in four capability domains: Event Grounding, Procedural Reasoning, Implicit Intent Inference, and Episodic Reasoning. The evaluation protocol separately measures video-to-plan reasoning, policy execution under oracle plans, and full task completion, in simulation and on a Franka Research 3 robot.

Released
2026-06-24
Readiness
Inspectable
Primary field
Robotics & Autonomous Systems

Why it matters

Existing manipulation benchmarks typically evaluate from a single current image, lacking grounding in observed human behavior. WatchAct fills this gap by assessing robots' ability to reason about events, procedures, intents, and scene changes from video, which is critical for real-world human-robot collaboration. It provides a disentangled evaluation to isolate reasoning and execution failures, offering practical value for developing and comparing systems.

Motivation

A robot working alongside people must reason about what they have done, in what order, and with what intent.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.