AI BENCHMARK PROFILE
HumanMoveVQA
HumanMoveVQA evaluates video MLLMs on reasoning about human trajectory and orientation changes in videos, using a first-frame anchored world coordinate system and 10K question-answer pairs across seven reasoning categories.
- Released
- 2026-06-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks fail to probe global human motion in space over time; HumanMoveVQA targets this gap, but without accessible data or code its practical value for model comparison remains unclear.
Motivation
Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motion into coarse semantic labels.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.