Benchmark Radar
AI BENCHMARK PROFILE

ForesightSafety-VLA

Robotics & Autonomous SystemsSafety & Trustworthiness

A diagnostic safety benchmark for vision-language-action models, evaluating policies across physical interaction, instruction, and perception safety with cumulative cost and risk exposure metrics.

Released
2026-06-25
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

The benchmark addresses the lack of systematic safety evaluation for embodied VLA models, offering process-level risk metrics and a taxonomy to localize failure sources, which supports model selection and safety improvements.

Motivation

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.