ExtremeWhenBench
ExtremeWhenBench evaluates natural-language temporal grounding in hour-long videos, with 2,273 open-form queries over 194 videos. Metrics include mIoU and R@k for predicted time intervals, with strict parse-failure handling.
- Released
- 2026-06-10
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides the first open hour-scale grounding benchmark to study the search problem in video understanding. Useful for developing and evaluating models for long-video retrieval and reasoning.
Motivation
Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-language grounding remain underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.