Benchmark Radar
AI BENCHMARK PROFILE

SciHazard

General AISafety & TrustworthinessDeharmScore GitHub

SciHazard evaluates LLMs on scientific safety risks across 12 disciplines with 2,400 hazardous and 600 oversafety questions grounded in regulated entities. The DeHarm-Score decomposes harm into executability and net-new risk, providing a detailed scoring contract.

Released
2026-07-21
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing safety benchmarks often use templated queries and LLM-as-a-Judge without domain grounding. SciHazard offers real-world grounded evaluation, and the DeHarm-Score improves agreement with expert annotations by 90% over baselines, enabling more reliable safety measurement for scientific LLMs and agents.

Motivation

Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.