Benchmark Radar
AI BENCHMARK PROFILE

GuardianBench

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across hazard categories, based on international safety standards. Evaluates VLM safety reasoning under latent contextual risk.

Released
2026-08-22
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Exposes a critical failure mode in embodied AI safety—instruction-insensitive verdicts—and provides a controlled suite for improvement research.

Motivation

In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.