Benchmark Radar
AI BENCHMARK PROFILE

NRT-Bench

General AISafety & Trustworthiness

Evaluates multi-turn red-teaming of LLM agents acting as operators of a simulated nuclear power plant control room. Safety is measured by objective loss of critical safety functions (CSFs) under adaptive attacks.

Released
2026-06-18
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a repeatable environment for assessing LLM agent robustness in safety-critical operations, where failure is objective rather than judged, and supports reproducible safety evaluation across models and defences.

Motivation

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.