SynthDocBench
SynthDocBench evaluates vision-language models on 1,788 questions over 200 synthetic long-context documents (avg. 51.1 pages) with 3,340 charts. It varies document length, layout archetype, modality composition, and question type as independent controlled factors, across chart, cross-modal, and complex subsets with deterministic ground truth.
- Released
- 2026-07-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing document benchmarks confound length, layout, and modality, obscuring specific model failure causes. SynthDocBench provides controlled attribution of failures, revealing sharp degradation with document length, positional sensitivity in the middle third, and breakdown of chart comprehension in long documents, which is valuable for targeted model improvement.
Motivation
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.