Benchmark Radar
AI BENCHMARK PROFILE

EgoSafe-Bench

General AISafety & Trustworthiness

EgoSafe-Bench evaluates visual safety understanding in first-person video, using 12,000 QA samples from 3,000 clips under the Hierarchical Reasoning Evaluation (HRE) protocol, which requires reasoning from feature anchoring to intent inference.

Released
2026-07-29
Readiness
Paper only
Primary field
General AI

Why it matters

Existing safety benchmarks rely on third-person footage and binary metrics, missing the causal reasoning gap in egocentric perception. EgoSafe-Bench provides a reusable protocol to assess whether LVLMs can move beyond correlation to forensic logic, informing model selection for safety-critical applications.

Motivation

Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.