Benchmark Radar
AI BENCHMARK PROFILE

MedHal-Loc

Health & Life SciencesKnowledge & Reasoning

MedHal-Loc is a benchmark and metric for localization faithfulness of medical hallucination detectors, comprising a controlled subset of 300 PubMedQA-derived statements with injected span-level errors and a natural subset. It evaluates whether top-ranked error units overlap erroneous spans across four error types.

Released
2026-06-19
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

The benchmark addresses the evaluation gap in measuring whether hallucination detectors that claim explainability actually localize errors faithfully, providing a way to assess detection and localization validity separately.

Motivation

Detecting hallucinations in clinical text is increasingly framed as an explainability problem: systems should not merely flag an unreliable response but point to the offending span.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.