Benchmark Radar
AI BENCHMARK PROFILE

EBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceShanghai AI Laboratory

EBench is a simulation benchmark for diagnosing generalist mobile manipulation policies. It comprises 26 manipulation tasks annotated along five capability dimensions (scene, atomic skill, horizon, precision, mobility) and four generalization dimensions (object, background, instruction, mixed), with strict train/test splits and held-out online evaluation.

Released
2026-06-16
Readiness
Runnable
Primary field
Robotics & Autonomous Systems

Why it matters

EBench addresses the need for multi-axis diagnostic evaluation of generalist manipulation models, moving beyond single success-rate metrics. It provides practical decision value by revealing capability and generalization profiles that guide model iteration and selection.

Motivation

We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.