MedStreamBench
MedStreamBench is a time-aware benchmark for medical video understanding, integrating 22 medical datasets and 5,419 QA instances across four temporal settings: retrospective, present, future, and proactive. Models are restricted to temporally bounded evidence windows and evaluated on answer correctness, responsiveness, and post-evidence stability.
- Released
- 2026-07-02
- Readiness
- Inspectable
- Primary field
- Health & Life Sciences
Why it matters
It addresses the gap between offline recognition and temporally grounded decision-making in clinical settings, where models must decide when to answer or alert. The benchmark provides a protocol for evaluating time-aware reasoning in medical video.
Motivation
Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.