Benchmark Radar
AI BENCHMARK PROFILE

KIDBench

General AISafety & Trustworthiness

KIDBench evaluates child-facing safety of LLMs for ages 7-11 using realistic queries and multi-turn child-actor simulations, scored by an LLM-as-a-Judge rubric.

Released
2026-05-25
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses a gap in LLM safety evaluation by focusing on age-appropriate responses for children, with a novel rubric and multi-turn assessment.

Motivation

Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidance, and boundaries.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.