SciHazard
SciHazard evaluates LLMs on scientific safety risks across 12 disciplines with 2,400 hazardous and 600 oversafety questions grounded in regulated entities. The DeHarm-Score decomposes harm into executability and net-new risk, providing a detailed scoring contract.
- Released
- 2026-07-21
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing safety benchmarks often use templated queries and LLM-as-a-Judge without domain grounding. SciHazard offers real-world grounded evaluation, and the DeHarm-Score improves agreement with expert annotations by 90% over baselines, enabling more reliable safety measurement for scientific LLMs and agents.
Motivation
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.