Benchmark Radar
AI BENCHMARK PROFILE

StreamArena

General AIMultimodal PerceptionJIA-Lab-research

StreamArena evaluates hour-scale streaming video understanding across 243 full-length videos (avg 88.8 min) with 3,646 open-ended QA pairs, covering real-time perception, historical retrospection, proactive interaction, and multimodal tool use. Includes a standardized runner and LLM-as-judge scorer.

Released
2026-08-06
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the lack of benchmarks for long-horizon, interactive streaming video understanding, where short clips and multiple-choice formats allow shortcuts. Provides a rigorous, open-ended evaluation to assess progress in continuous, interactive multimodal agents.

Motivation

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.