Benchmark Radar
AI BENCHMARK PROFILE

GST-Bench

Robotics & Autonomous SystemsMultimodal Perception

GST-Bench is a VQA benchmark for global spatial intelligence in video understanding, covering synthetic videos and human-verified questions. It evaluates VLMs' ability to infer spatial relations from novel viewpoints and map egocentric observations to top-down views.

Released
2026-08-06
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Existing video benchmarks focus on local spatial perception, while GST-Bench targets global spatial awareness over long-horizon videos, addressing a gap in evaluating embodied agents' spatial intelligence. It provides a scoring contract to compare VLMs and highlights a significant performance gap between models and humans.

Motivation

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.