Benchmark Radar
AI BENCHMARK PROFILE

ManiGuard-Bench

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

ManiGuard-Bench evaluates 1,000 locked scenarios across six contact-rich household task families, with safety specifications checked by LTL$_f$-grounded automaton monitors and metrics separating safe success from engaged-and-safe behavior.

Released
2026-08-18
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

ManiGuard-Bench fills a gap by evaluating robotic manipulation safety independently of task success, revealing that up to 21% of successful rollouts violate safety specifications and enabling safety-aware fine-tuning with 8,000 demonstrations.

Motivation

Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.