Benchmark Radar
AI BENCHMARK PROFILE

HalluTruthQA

General AIMultimodal Perception

HalluTruthQA is a fine-grained benchmark for Arabic QA hallucination detection, localization, and explanation, containing 2,400 expert-curated examples across four knowledge domains with span-level and explanation annotations.

Released
2026-07-22
Readiness
Paper only
Primary field
General AI

Why it matters

Hallucination evaluation typically uses response-level labels; this benchmark provides granular annotations to assess localization and explanation capabilities, but lacks a public reuse path for broader comparison.

Motivation

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.