Benchmark Radar
AI BENCHMARK PROFILE

EgoSafetyBench

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

EgoSafetyBench is a diagnostic egocentric video benchmark of 1,200 robot-view scenarios to evaluate vision-language models as runtime safety guards. It assesses situational awareness across routine, suspicious, obvious, and contextual hazards, and visual-channel robustness against misleading in-scene text.

Released
2026-06-30
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

This benchmark addresses the practical need for safety guards that distinguish genuine hazards from superficially alarming but benign actions, and highlights the vulnerability to misleading signage, which can inform safer deployment of embodied VLMs in real-world settings.

Motivation

Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.