Benchmark Radar
AI BENCHMARK PROFILE

ViSTR-Bench

General AIMultimodal Perception

An evaluation suite of 1,340 video QA pairs across 15 subtasks assessing MLLM qualitative reasoning in dynamic scenes, covering motion perception, spatial relations, outcome prediction, and physical dynamics.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

Current MLLMs lag behind humans in intuitive spatial-temporal reasoning; this probe highlights those gaps but lacks a standalone reusable benchmark.

Motivation

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic reasoning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.