AI BENCHMARK PROFILE
ForesightSafety-VLA
A diagnostic safety benchmark for vision-language-action models, evaluating policies across physical interaction, instruction, and perception safety with cumulative cost and risk exposure metrics.
- Released
- 2026-06-25
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
The benchmark addresses the lack of systematic safety evaluation for embodied VLA models, offering process-level risk metrics and a taxonomy to localize failure sources, which supports model selection and safety improvements.
Motivation
In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.