Benchmark Radar
AI BENCHMARK PROFILE

ExtractBench

General AIKnowledge & ReasoningLlamaIndex

ExtractBench evaluates schema-guided extraction from enterprise documents. Given a document and a user-defined JSON schema, systems must return schema-valid JSON with correct values, include every record of repeated structures, mark missing fields as null, and provide source evidence. The benchmark includes 4,869 pages across 370 documents, 8 business domains, and 67 document types, with tags for challenge, perception, table structure, length, and domain. Scoring uses unified value F1 for value accuracy and word- and page-level F1 for grounding.

Released
2026-07-31
Readiness
Runnable
Primary field
General AI

Why it matters

Enterprise workflows increasingly rely on agents for schema-guided extraction, where errors can lead to wrong payments or decisions. ExtractBench addresses the lack of a benchmark that jointly measures value accuracy, completeness, grounding, and cost, providing a standardized evaluation for comparing extraction systems on realistic document types and lengths.

Motivation

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.