Benchmark Radar
AI BENCHMARK PROFILE

LongDocBench

General AIKnowledge & ReasoningLongDocBench Team

LongDocBench is a benchmark for Table-of-Contents Hierarchy Recovery and Contextual Relationship Recovery in long documents. It includes 85 real-world documents (financial reports, textbooks, academic papers) spanning 2,582 pages, with human-verified annotations for 3,937 heading nodes and 3,258 contextual relationships.

Released
2026-08-15
Readiness
Paper only
Primary field
General AI

Why it matters

Existing document parsing benchmarks focus on page-level tasks, leaving document-level structure recovery unevaluated. LongDocBench provides a standardized evaluation for tasks that are critical for understanding long documents, enabling comparison of parsers on hierarchy and relationship recovery.

Motivation

Parsing visual documents into machine-readable representations is fundamental to document intelligence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.