Benchmark Radar
AI BENCHMARK PROFILE

SafeBuild-Bench

General AISafety & TrustworthinessSafeBuild

A benchmark for evaluating multimodal large language models on construction safety hazard identification and description, with 3,314 expert-verified task instances from over 3,000 images.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

Construction-safety models must handle realistic temporal and site variation; this benchmark provides a standard evaluation for hazard recognition and description, useful for deployment risk assessment.

Motivation

Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only recognize common objects in curated images.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.