AI BENCHMARK PROFILE
GuardianBench
3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across hazard categories, based on international safety standards. Evaluates VLM safety reasoning under latent contextual risk.
- Released
- 2026-08-22
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Exposes a critical failure mode in embodied AI safety—instruction-insensitive verdicts—and provides a controlled suite for improvement research.
Motivation
In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.