CrossView
CrossView is a multi-camera video question-answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. It evaluates vision-language models on tasks that require joint reasoning across multiple simultaneous camera views, including resolving occlusions and integrating evidence across perspectives.
- Released
- 2026-08-16
- Readiness
- Inspectable
- Primary field
- Transport & Logistics
Why it matters
Existing video benchmarks focus on single-camera settings, leaving multi-camera reasoning unmeasured. CrossView provides a standardized evaluation for a capability critical to real-world applications like autonomous vehicles and surveillance, enabling comparison of models on tasks that scale with viewpoint number and require cross-view integration.
Motivation
Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.