Benchmark Radar
AI BENCHMARK PROFILE

WorldRoamBench

General AIMultimodal Perception

WorldRoamBench evaluates interactive world models on long-horizon stability across action, vision, physics, and memory dimensions. It includes 600+ test cases in various scenes and views with continuous interaction.

Released
2026-06-30
Readiness
Paper only
Primary field
General AI

Why it matters

Interactive world models need stable, physically grounded, and memory-faithful behavior, but existing benchmarks ignore these aspects. WorldRoamBench provides a comprehensive evaluation revealing that no current model reliably satisfies all dimensions.

Motivation

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.