Benchmark Radar
AI BENCHMARK PROFILE

SteelBench

General AISafety & Trustworthiness

STEELBENCH evaluates vision-language models on per-worker activity recognition and safety-rule reasoning in industrial CCTV footage. It includes 1,345 clips with dense annotations and a provenance-aware audit protocol.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

Industrial surveillance poses unique visual and procedural challenges. This benchmark reveals large gaps in VLM performance and shows how annotation provenance can inflate accuracy, guiding reliable deployment.

Motivation

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.