ViSTR-Bench
An evaluation suite of 1,340 video QA pairs across 15 subtasks assessing MLLM qualitative reasoning in dynamic scenes, covering motion perception, spatial relations, outcome prediction, and physical dynamics.
- Released
- 2026-07-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current MLLMs lag behind humans in intuitive spatial-temporal reasoning; this probe highlights those gaps but lacks a standalone reusable benchmark.
Motivation
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic reasoning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.