ManiGuard-Bench
ManiGuard-Bench evaluates 1,000 locked scenarios across six contact-rich household task families, with safety specifications checked by LTL$_f$-grounded automaton monitors and metrics separating safe success from engaged-and-safe behavior.
- Released
- 2026-08-18
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
ManiGuard-Bench fills a gap by evaluating robotic manipulation safety independently of task success, revealing that up to 21% of successful rollouts violate safety specifications and enabling safety-aware fine-tuning with 8,000 demonstrations.
Motivation
Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.