AI BENCHMARK PROFILE
OVO-S-Bench
OVO-S-Bench evaluates streaming spatial intelligence in MLLMs with 1,680 human-annotated questions across four abstraction levels, using streaming prefixes and evidence intervals.
- Released
- 2026-06-02
- Readiness
- Runnable
- Primary field
- Transport & Logistics
Why it matters
Addresses the gap in evaluating spatial reasoning from continuous egocentric streams, providing a standardized benchmark with human annotation and clear evaluation protocols.
Motivation
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.