Benchmark Radar
AI BENCHMARK PROFILE

OVO-S-Bench

Transport & LogisticsMultimodal PerceptionInternLM

OVO-S-Bench evaluates streaming spatial intelligence in MLLMs with 1,680 human-annotated questions across four abstraction levels, using streaming prefixes and evidence intervals.

Released
2026-06-02
Readiness
Runnable
Primary field
Transport & Logistics

Why it matters

Addresses the gap in evaluating spatial reasoning from continuous egocentric streams, providing a standardized benchmark with human annotation and clear evaluation protocols.

Motivation

Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.