MiraBench
MiraBench evaluates action-conditioned reliability in robotic world models through three levels: Physics Adherence, Action-Following Fidelity, and Optimism Bias Detection, using a human-annotated corpus of over 16,000 judgments.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing benchmarks focus on visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to actions, and calibrated to failure. MiraBench aims to provide a diagnostic foundation for assessing world models as simulators, but its public evaluation protocol and release details are not yet specified.
Motivation
Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.