AI BENCHMARK PROFILE
XL-DocBench
XL-DocBench evaluates evidence-grounded long-document understanding with 1,519 human-verified questions from six professional domains, contexts up to 2,303 pages, multi-page evidence, and typed reasoning rules.
- Released
- 2026-07-21
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
Professional workflows require traceable answers from long documents; this benchmark fills a gap in multi-page and structured reasoning evaluation, enabling failure attribution.
Motivation
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.