Benchmark Radar
AI BENCHMARK PROFILE

Doc2DB-Bench

General AIKnowledge & ReasoningDoc2DB-Bench Team

Doc2DB-Bench evaluates document-to-database construction, converting long heterogeneous documents into normalized relational databases with entity identities, keys, cross-table links, and integrity constraints. It includes 203 document instances across 42 schemas and 7 domains, with fine-grained capability annotations for intra-table extraction and inter-table reasoning.

Released
2026-08-09
Readiness
Runnable
Primary field
General AI

Why it matters

Existing document-to-table benchmarks overlook relational database requirements such as normalization, entity resolution, and cross-table consistency. Doc2DB-Bench addresses this gap by providing a testbed for assessing LLM-based systems in realistic database construction, supporting analytics, compliance, and decision-making applications.

Motivation

Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.