WBench
WBench evaluates interactive video world models across five dimensions: video quality, setting adherence, interaction adherence, consistency, and physics compliance. It includes 289 test cases and 1,058 multi-turn interaction sequences, covering diverse scenes and control types, with 22 automatic sub-metrics validated against human judgment.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
No unified standard previously existed for evaluating interactive world models across the required competencies. WBench provides a comprehensive, multi-turn benchmark with a public leaderboard and open data/code, enabling systematic model comparison and diagnostic insights into strengths and weaknesses.
Motivation
Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.