Benchmark Radar
AI BENCHMARK PROFILE

ConfBench

General AIMultimodal Perception

ConfBench is a calibration-specific benchmark for key information extraction from documents. It applies 20 degradation pipelines to create 1,346 variants and over 70K entity-level evaluations, spanning the accuracy spectrum for evaluating confidence estimates of VLMs.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Document processing requires trustworthy confidence scores for routing automation vs. human review. ConfBench enables systematic study of confidence estimators and calibration methods, addressing the lack of calibration-focused benchmarks.

Motivation

Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.