VGenST-Bench
VGenST-Bench evaluates spatio-temporal reasoning in multimodal large language models using 1,200 procedurally generated videos with controlled scene composition, camera trajectory, and reasoning targets. It covers 12 reasoning tasks across three spatial scales and 12 QA types across three reasoning levels, with multiple choice and open-ended question variants.
- Released
- 2026-05-21
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing video reasoning benchmarks rely on static or passively curated content, limiting fine-grained diagnosis. VGenST-Bench uses active synthesis to enable controlled and diverse evaluation of fine-grained spatio-temporal reasoning, supporting model comparison and targeted improvement in MLLMs.
Motivation
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.