Benchmark Radar
AI BENCHMARK PROFILE

SynthDocBench

General AIMultimodal PerceptionLong Context & MemoryServiceNow AI

SynthDocBench evaluates vision-language models on 1,788 questions over 200 synthetic long-context documents (avg. 51.1 pages) with 3,340 charts. It varies document length, layout archetype, modality composition, and question type as independent controlled factors, across chart, cross-modal, and complex subsets with deterministic ground truth.

Released
2026-07-11
Readiness
Runnable
Primary field
General AI

Why it matters

Existing document benchmarks confound length, layout, and modality, obscuring specific model failure causes. SynthDocBench provides controlled attribution of failures, revealing sharp degradation with document length, positional sensitivity in the middle third, and breakdown of chart comprehension in long documents, which is valuable for targeted model improvement.

Motivation

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.