SafeRelBench
SafeRelBench evaluates VLM-driven embodied agents on household tasks with process-level safety constraints, focusing on spatial relations such as support, containment, and proximity. It includes 507 executable samples (248 spatial-relation, 259 control) and measures whether agents satisfy safety conditions before risk-prone actions, alongside task success.
- Released
- 2026-07-16
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Safety in embodied agents depends on spatial awareness during action sequences, not just final outcomes. SafeRelBench fills a gap by quantifying process-level safety compliance, enabling comparison across agents and highlighting the need for improved spatial reasoning in safe planning.
Motivation
Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.