Benchmark Radar
AI BENCHMARK PROFILE

CrossView

Transport & LogisticsMultimodal PerceptionUT Austin Swarm Lab

CrossView is a multi-camera video question-answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. It evaluates vision-language models on tasks that require joint reasoning across multiple simultaneous camera views, including resolving occlusions and integrating evidence across perspectives.

Released
2026-08-16
Readiness
Inspectable
Primary field
Transport & Logistics

Why it matters

Existing video benchmarks focus on single-camera settings, leaving multi-camera reasoning unmeasured. CrossView provides a standardized evaluation for a capability critical to real-world applications like autonomous vehicles and surveillance, enabling comparison of models on tasks that scale with viewpoint number and require cross-view integration.

Motivation

Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.