Benchmark Radar
AI BENCHMARK PROFILE

MBench

General AIMultimodal Perception

MBench is a benchmark for memory capability of video world models, decomposing into entity, environment, and causal consistency with 12 sub-dimensions. Includes code, dataset, and leaderboard.

Released
2026-05-30
Readiness
Runnable
Primary field
General AI

Why it matters

Fills the gap in evaluating long-term state retention in video world models, providing a standardized benchmark to advance the field.

Motivation

Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.