Benchmark Radar
AI BENCHMARK PROFILE

RealDocBench

General AIMultimodal Perception

RealDocBench evaluates field-level QA and layout understanding on real regulated documents, with 1,356 field-level questions over 581 documents and 1,500 annotated page images, scored on per-field accuracy and adjacency-aware layout metrics.

Released
2026-06-05
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the gap in document parsing evaluation by focusing on real-world regulated documents and specific field-level needs, enabling cost-aware comparisons of commercial and open-source systems.

Motivation

Document parsing systems are increasingly deployed in high-stakes, regulated workflows such as mortgage underwriting, financial reporting, supply-chain logistics, and clinical records.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.