Benchmark Radar
AI BENCHMARK PROFILE

SafeRelBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceAuthors of SafeRelBench

SafeRelBench evaluates VLM-driven embodied agents on household tasks with process-level safety constraints, focusing on spatial relations such as support, containment, and proximity. It includes 507 executable samples (248 spatial-relation, 259 control) and measures whether agents satisfy safety conditions before risk-prone actions, alongside task success.

Released
2026-07-16
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Safety in embodied agents depends on spatial awareness during action sequences, not just final outcomes. SafeRelBench fills a gap by quantifying process-level safety compliance, enabling comparison across agents and highlighting the need for improved spatial reasoning in safe planning.

Motivation

Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.