AI BENCHMARK PROFILE
KIDBench
KIDBench evaluates child-facing safety of LLMs for ages 7-11 using realistic queries and multi-turn child-actor simulations, scored by an LLM-as-a-Judge rubric.
- Released
- 2026-05-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses a gap in LLM safety evaluation by focusing on age-appropriate responses for children, with a novel rubric and multi-turn assessment.
Motivation
Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidance, and boundaries.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.