Benchmark Radar
AI BENCHMARK PROFILE

AuthorityBench

General AISafety & Trustworthinessfloating-reeds

AuthorityBench evaluates how citation-based authority signals affect epistemic behavior in LLMs using 220,564 prompts across four domains, with a 2x2 factorial design crossing claim veracity and citation veracity.

Released
2026-06-11
Readiness
Runnable
Primary field
General AI

Why it matters

The benchmark isolates citation presence from content to measure susceptibility to citation-induced hallucination, providing a standardized protocol for assessing reliability in citation-augmented settings.

Motivation

Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.