MIRA-Ev
MIRA-Ev evaluates clinical NLP on evidence detection and relational reasoning in Spanish MIR exam cases, with three tasks: evidence sentence retrieval, argumentative component extraction, and relation classification.
- Released
- 2026-07-21
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
It addresses the lack of evidence-level evaluation in clinical NLP, providing a multilingual benchmark (Spanish, English, Basque) to assess not just final answers but the grounding of diagnoses.
Motivation
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.