Benchmark Radar
AI BENCHMARK PROFILE

ReceiptBench

General AIMultimodal PerceptionT0rl (ReceiptBench team)

ReceiptBench evaluates multimodal large language models on visual information extraction from receipts. It includes 10k human-annotated receipts and four hierarchical subtasks: basic perception, format normalization, semantic reasoning, and structure parsing, with defined metrics and scoring.

Released
2026-05-21
Readiness
Runnable
Primary field
General AI

Why it matters

Existing VIE benchmarks lack scale, realism, and semantic granularity. ReceiptBench provides a standardized, publicly available evaluation path for comparing models on diverse receipt understanding tasks, supporting practical deployment decisions in document automation.

Motivation

Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.