OmniInteract
OmniInteract is a streaming benchmark for real-time omnimodal LLMs evaluated through native online inference over audio-visual streams. It contains 250 videos with 1,430 temporally grounded response slots (1Q1A and 1QnA), with each slot including trigger, response window, and target answer. Metrics include IA-QTF1, Interruption Diagnostic Suite, and Nested Chain Completion Score.
- Released
- 2026-05-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Real-time omnimodal assistants must process streaming audio-visual input and decide whether and when to respond without access to future content, a capability not captured by offline video benchmarks. OmniInteract provides a native streaming evaluation protocol, revealing that current models remain weak in streaming interaction.
Motivation
We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.