AI BENCHMARK PROFILE
SafeGesture
SafeGesture evaluates vision-language models on scenario-conditioned safety interpretation of hand gestures, pairing 6 gestures with 8 scenarios for 4,800 items.
- Released
- 2026-08-17
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It exposes a perception-reasoning gap in safety-critical gesture interpretation, providing a reusable test for scenario-dependent decision-making.
Motivation
Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained hand gestures in safety-critical operational contexts remains largely unexamined.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.