CVSBench
CVSBench evaluates cross-view spatial reasoning in vision-language models using satellite-street image pairs. It includes tasks for cross-view VQA, grounding, and viewpoint identification, with 3,297 image groups, 9,468 object-level annotations, and 40,679 QA pairs.
- Released
- 2026-06-21
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
CVSBench addresses the gap in evaluating VLMs' ability to reason about scenes across drastically different viewpoints, which is crucial for applications like navigation and remote sensing. It provides a systematic protocol for measuring object-level and layout consistency under viewpoint changes.
Motivation
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.