AI BENCHMARK PROFILE
EgoSAT
EgoSAT evaluates vision-language models on egocentric video reasoning in streaming settings. It contains 1,997 videos (165 hours) and about 4,800 QA pairs covering retrospective, online, and prospective reasoning tasks.
- Released
- 2026-06-23
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Serves as a unified benchmark for streaming egocentric interaction understanding, enabling assessment of temporal reasoning and confidence calibration in VLMs, which is crucial for reliable deployment in real-time applications.
Motivation
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models (VLMs).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.