AI BENCHMARK PROFILE
WorldRoamBench
WorldRoamBench evaluates interactive world models on long-horizon stability across action, vision, physics, and memory dimensions. It includes 600+ test cases in various scenes and views with continuous interaction.
- Released
- 2026-06-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Interactive world models need stable, physically grounded, and memory-faithful behavior, but existing benchmarks ignore these aspects. WorldRoamBench provides a comprehensive evaluation revealing that no current model reliably satisfies all dimensions.
Motivation
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.