Benchmark Radar
AI BENCHMARK PROFILE

MiraBench

Robotics & Autonomous SystemsKnowledge & Reasoning

MiraBench evaluates action-conditioned reliability in robotic world models through three levels: Physics Adherence, Action-Following Fidelity, and Optimism Bias Detection, using a human-annotated corpus of over 16,000 judgments.

Released
2026-05-28
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Existing benchmarks focus on visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to actions, and calibrated to failure. MiraBench aims to provide a diagnostic foundation for assessing world models as simulators, but its public evaluation protocol and release details are not yet specified.

Motivation

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.