Benchmark Radar
AI BENCHMARK PROFILE

CAREBench

General AISafety & Trustworthiness

CAREBench evaluates language models on upstream child-safety risks with 500 prompts across twelve categories, assessing recognition, refusal, de-escalation, and redirection.

Released
2026-06-29
Readiness
Paper only
Primary field
General AI

Why it matters

Child-safety evaluation often focuses on explicit material; CAREBench targets earlier risk scenarios, which could help developers identify policy gaps.

Motivation

How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.