Benchmark Radar
AI BENCHMARK PROFILE

LongEgoRefer

General AIMultimodal Perception

LongEgoRefer is a benchmark for Video Referring Expression Comprehension in long-form egocentric videos, built from Ego4D. It contains 1,498 referring expressions over videos averaging 45 minutes, requiring temporal grounding of when an occurrence happens and spatial grounding of where the object appears. The benchmark uses evaluation metrics for temporal and spatial grounding.

Released
2026-07-02
Readiness
Runnable
Primary field
General AI

Why it matters

Existing egocentric Video REC benchmarks focus on short clips, not reflecting real-world long-form recordings. This benchmark defines a demanding spatio-temporal grounding problem that tests models on sparse object occurrences and complex interactions over extended sequences.

Motivation

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.