ATOM-Bench
ATOM-Bench evaluates atomic skills and compositional generalization in manipulation policies across 30 atomic tasks and 24 held-out compositional tasks, using paired single-arm and dual-arm robot tracks, with 3,000 human demonstrations and evaluation rollout data released.
- Released
- 2026-06-15
- Readiness
- Inspectable
- Primary field
- Robotics & Autonomous Systems
Why it matters
ATOM-Bench provides a public diagnostic testbed for disentangling failures in motor execution, instruction grounding, and compositional reuse, which is crucial for advancing generalist manipulation policies beyond demonstrated tasks.
Motivation
Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization remains difficult to diagnose.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.