Benchmark Radar
AI BENCHMARK PROFILE

MIRA-Ev

Health & Life SciencesMultimodal Perception

MIRA-Ev evaluates clinical NLP on evidence detection and relational reasoning in Spanish MIR exam cases, with three tasks: evidence sentence retrieval, argumentative component extraction, and relation classification.

Released
2026-07-21
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

It addresses the lack of evidence-level evaluation in clinical NLP, providing a multilingual benchmark (Spanish, English, Basque) to assess not just final answers but the grounding of diagnoses.

Motivation

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.