Benchmark Radar
AI BENCHMARK PROFILE

OpenHalDet

CybersecurityMultimodal Perception

OpenHalDet is a unified benchmark for hallucination detection across 17 datasets, supporting black-box, gray-box, and white-box detectors with standardized pipelines, scoring via AUROC and Cost@N.

Released
2026-06-05
Readiness
Runnable
Primary field
Cybersecurity

Why it matters

It standardizes hallucination detection evaluation, enabling fair comparison across diverse methods and providing a systematic view of detector performance in LLM applications, addressing inconsistencies and limited coverage.

Motivation

Hallucination detection is essential for the reliable deployment of large language models (LLMs).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.