AI BENCHMARK PROFILE
CAREBench
CAREBench evaluates language models on upstream child-safety risks with 500 prompts across twelve categories, assessing recognition, refusal, de-escalation, and redirection.
- Released
- 2026-06-29
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Child-safety evaluation often focuses on explicit material; CAREBench targets earlier risk scenarios, which could help developers identify policy gaps.
Motivation
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.