Benchmark Radar
AI BENCHMARK PROFILE

XL-DocBench

Health & Life SciencesFinance & EconomicsKnowledge & Reasoning

XL-DocBench evaluates evidence-grounded long-document understanding with 1,519 human-verified questions from six professional domains, contexts up to 2,303 pages, multi-page evidence, and typed reasoning rules.

Released
2026-07-21
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Professional workflows require traceable answers from long documents; this benchmark fills a gap in multi-page and structured reasoning evaluation, enabling failure attribution.

Motivation

Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.