AI BENCHMARK PROFILE
HalluTruthQA
HalluTruthQA is a fine-grained benchmark for Arabic QA hallucination detection, localization, and explanation, containing 2,400 expert-curated examples across four knowledge domains with span-level and explanation annotations.
- Released
- 2026-07-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Hallucination evaluation typically uses response-level labels; this benchmark provides granular annotations to assess localization and explanation capabilities, but lacks a public reuse path for broader comparison.
Motivation
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.