Benchmark Radar
AI BENCHMARK PROFILE

SafeGesture

General AISafety & TrustworthinessThe Responsible AI Initiative

SafeGesture evaluates vision-language models on scenario-conditioned safety interpretation of hand gestures, pairing 6 gestures with 8 scenarios for 4,800 items.

Released
2026-08-17
Readiness
Runnable
Primary field
General AI

Why it matters

It exposes a perception-reasoning gap in safety-critical gesture interpretation, providing a reusable test for scenario-dependent decision-making.

Motivation

Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained hand gestures in safety-critical operational contexts remains largely unexamined.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.