Benchmark Radar
AI BENCHMARK PROFILE

HealthBench-Psych

Health & Life SciencesKnowledge & Reasoning

Evaluates LLMs on 610 mental-health conversations from HealthBench using a three-judge panel and physician rubrics, with a harder 119-conversation subset.

Released
2026-08-25
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Provides a clinically validated mental-health evaluation subset with released data, code, and judge panel, enabling reproducible specialty-specific model assessment.

Motivation

General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.