Benchmark Radar
AI BENCHMARK PROFILE

MedDDC-Eval

Health & Life SciencesKnowledge & Reasoning

MedDDC-Eval evaluates multi-turn medical consultation agents by decoupling diagnosis from the diagnostic reader, using a frozen shared reader to score policies. It reports diagnostic support, coverage, and efficiency across Record and Dialogue splits.

Released
2026-07-21
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Coupled evaluation confounds history elicitation with terminal diagnosis. This testbed enables fair comparison and evaluation-informed policy optimization, providing a more reliable measure of diagnostic support.

Motivation

Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.