Benchmark Radar
AI BENCHMARK PROFILE

ArmnetBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceArmnetBench Team

A benchmark for robot manipulation policies evaluated on a fleet of low-cost SO-101 cells. It compares 7 policies across 12 tasks in single-arm and bimanual configurations, with 2,518 policy rollouts and 600 reference demonstrations, all labeled successful, suboptimal, or failure. Data is released in LeRobot and RoboMeter formats.

Released
2026-07-27
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Real-world evaluation of manipulation policies is costly and difficult to standardize. This benchmark provides a shared, public protocol with quality-labeled data, enabling comparable assessment of policies and supporting research on learning from mixed-quality demonstrations.

Motivation

Real-world evaluation is a bottleneck in developing generalist robot manipulation policies.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.