Benchmark Radar
AI BENCHMARK PROFILE

EMPATH

General AISafety & Trustworthiness

EMPATH evaluates safety of emotional-support chatbots via auditor-generated multi-turn conversations scored on 19 metrics across crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation.

Released
2026-06-29
Readiness
Paper only
Primary field
General AI

Why it matters

Safety evaluation for emotional-support chatbots needs multilingual, multi-turn metrics; EMPATH appears to address that gap.

Motivation

Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.