EgoSafetyBench
EgoSafetyBench is a diagnostic egocentric video benchmark of 1,200 robot-view scenarios to evaluate vision-language models as runtime safety guards. It assesses situational awareness across routine, suspicious, obvious, and contextual hazards, and visual-channel robustness against misleading in-scene text.
- Released
- 2026-06-30
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
This benchmark addresses the practical need for safety guards that distinguish genuine hazards from superficially alarming but benign actions, and highlights the vulnerability to misleading signage, which can inform safer deployment of embodied VLMs in real-world settings.
Motivation
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.