Benchmark Radar
AI BENCHMARK PROFILE

MamaBench

Health & Life SciencesSafety & Trustworthiness

MamaBench evaluates LLM robustness in maternal and child health diagnosis using counterfactual clinical narratives, with the Bias Trap Rate (BTR) metric.

Released
2026-07-15
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

It highlights the gap between base accuracy and robust accuracy in clinical AI, showing that models can fail on clinically similar cases requiring different interventions.

Motivation

Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.