Benchmark Radar
AI BENCHMARK PROFILE

RoboTrustBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceRoboTrustBench Team

RoboTrustBench evaluates the trustworthiness of video world models for robotic manipulation across four scenarios: Normal, Constraint-Sensitive, Counterfactual, and Adversarial. It contains 1,207 expert-validated instruction-image pairs from DROID episodes and a six-dimensional evaluation protocol with 13 fine-grained criteria.

Released
2026-06-01
Readiness
Inspectable
Primary field
Robotics & Autonomous Systems

Why it matters

Existing benchmarks for video world models largely overlook trustworthiness aspects such as constraint reasoning, counterfactual grounding, and safety. RoboTrustBench provides a structured evaluation to assess these capabilities, offering practical guidance for selecting and improving models for safe robotic manipulation.

Motivation

Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.