AI BENCHMARK PROFILE
CRA-Bench
A session-layer framework for tracking conversational risk accumulation in multi-turn LLM systems, including CRA-Bench datasets and trajectory-native evaluation protocols.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating guardrails for multi-turn dialogues where benign turns compose into harm, providing session-level scoring metrics beyond isolated prompt-response checks.
Motivation
Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.