Benchmark Radar
AI BENCHMARK PROFILE

X-Stream

Transport & LogisticsMultimodal PerceptionMMLab, CUHKHuawei Inc.

Evaluates multimodal large language models on multi-stream streaming understanding, with 4,220 QA pairs across 932 videos covering 11 subtasks in multi-window, multi-view, and multi-device scenarios, using a dual-verification construction pipeline and online inference under a fixed average video-token rate.

Released
2026-06-01
Readiness
Runnable
Primary field
Transport & Logistics

Why it matters

Existing benchmarks focus on single-stream video understanding, leaving a gap for evaluating concurrent video streams common in live sports, autonomous driving, and multi-screen applications. This benchmark provides a practical evaluation protocol for multi-stream reasoning and exposes limitations in current models.

Motivation

While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inherently demand continuous, multi-stream interactions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.