Benchmark Radar
AI BENCHMARK PROFILE

EgoSAT

General AIMultimodal PerceptionEgoSAT Project

EgoSAT evaluates vision-language models on egocentric video reasoning in streaming settings. It contains 1,997 videos (165 hours) and about 4,800 QA pairs covering retrospective, online, and prospective reasoning tasks.

Released
2026-06-23
Readiness
Inspectable
Primary field
General AI

Why it matters

Serves as a unified benchmark for streaming egocentric interaction understanding, enabling assessment of temporal reasoning and confidence calibration in VLMs, which is crucial for reliable deployment in real-time applications.

Motivation

We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models (VLMs).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.