AI BENCHMARK PROFILE
CardioLens
CardioLens evaluates MLLMs on multi-sequence cardiac MRI interpretation, covering image understanding, report generation, and disease diagnosis using QA pairs from private hospital archives.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
The benchmark highlights the gap between public medical benchmark performance and clinical use, but its private data and lack of a public reuse path limit its value for broader model comparison.
Motivation
Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.