AI BENCHMARK PROFILE
MamaBench
MamaBench evaluates LLM robustness in maternal and child health diagnosis using counterfactual clinical narratives, with the Bias Trap Rate (BTR) metric.
- Released
- 2026-07-15
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
It highlights the gap between base accuracy and robust accuracy in clinical AI, showing that models can fail on clinically similar cases requiring different interventions.
Motivation
Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.