ArmnetBench
A benchmark for robot manipulation policies evaluated on a fleet of low-cost SO-101 cells. It compares 7 policies across 12 tasks in single-arm and bimanual configurations, with 2,518 policy rollouts and 600 reference demonstrations, all labeled successful, suboptimal, or failure. Data is released in LeRobot and RoboMeter formats.
- Released
- 2026-07-27
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Real-world evaluation of manipulation policies is costly and difficult to standardize. This benchmark provides a shared, public protocol with quality-labeled data, enabling comparable assessment of policies and supporting research on learning from mixed-quality demonstrations.
Motivation
Real-world evaluation is a bottleneck in developing generalist robot manipulation policies.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.