LongEgoRefer
LongEgoRefer is a benchmark for Video Referring Expression Comprehension in long-form egocentric videos, built from Ego4D. It contains 1,498 referring expressions over videos averaging 45 minutes, requiring temporal grounding of when an occurrence happens and spatial grounding of where the object appears. The benchmark uses evaluation metrics for temporal and spatial grounding.
- Released
- 2026-07-02
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing egocentric Video REC benchmarks focus on short clips, not reflecting real-world long-form recordings. This benchmark defines a demanding spatio-temporal grounding problem that tests models on sparse object occurrences and complex interactions over extended sequences.
Motivation
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.