AI BENCHMARK PROFILE
MedDDC-Eval
MedDDC-Eval evaluates multi-turn medical consultation agents by decoupling diagnosis from the diagnostic reader, using a frozen shared reader to score policies. It reports diagnostic support, coverage, and efficiency across Record and Dialogue splits.
- Released
- 2026-07-21
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Coupled evaluation confounds history elicitation with terminal diagnosis. This testbed enables fair comparison and evaluation-informed policy optimization, providing a more reliable measure of diagnostic support.
Motivation
Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.