AI BENCHMARK PROFILE
OpenHalDet
OpenHalDet is a unified benchmark for hallucination detection across 17 datasets, supporting black-box, gray-box, and white-box detectors with standardized pipelines, scoring via AUROC and Cost@N.
- Released
- 2026-06-05
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
It standardizes hallucination detection evaluation, enabling fair comparison across diverse methods and providing a systematic view of detector performance in LLM applications, addressing inconsistencies and limited coverage.
Motivation
Hallucination detection is essential for the reliable deployment of large language models (LLMs).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.