Benchmark Radar
AI BENCHMARK PROFILE

ExtremeWhenBench

General AIMultimodal PerceptionNAVER AI

ExtremeWhenBench evaluates natural-language temporal grounding in hour-long videos, with 2,273 open-form queries over 194 videos. Metrics include mIoU and R@k for predicted time intervals, with strict parse-failure handling.

Released
2026-06-10
Readiness
Runnable
Primary field
General AI

Why it matters

Provides the first open hour-scale grounding benchmark to study the search problem in video understanding. Useful for developing and evaluating models for long-video retrieval and reasoning.

Motivation

Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-language grounding remain underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.