Benchmark Radar
AI BENCHMARK PROFILE

LegalCiteTrust

General AIKnowledge & Reasoning

LegalCiteTrust evaluates citation trustworthiness in Chinese long-form legal research reports, assessing coverage, support, and citation-level existence, fidelity, and applicability.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

Long-form legal research reports increasingly rely on LLMs, but citation trustworthiness is critical for legal accuracy. This benchmark addresses the gap by measuring whether citations are not only real but also accurate and applicable, providing a more nuanced evaluation than simple existence checks.

Motivation

Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, but also on whether cited legal authorities are trustworthy.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.