AI BENCHMARK PROFILE
MMed-Bench-IR
MMed-Bench-IR evaluates multilingual medical information retrieval across 6 languages with three tasks: cross-lingual QA retrieval, concept discrimination, and evidence retrieval for RAG.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Uncovers severe cross-lingual failures in biomedical encoders that English-only benchmarks miss, highlighting the need for multilingual capability measurement in clinical RAG.
Motivation
Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.