Benchmark Radar
AI BENCHMARK PROFILE

OmniHandwritingOCR

General AIMultimodal Perception

Evaluates multimodal LLMs and OCR systems on handwritten text and mathematical expression recognition across six subtasks and twelve subsets with 77.57K images.

Released
2026-08-19
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a challenging diagnostic benchmark for realistic handwritten OCR, highlighting failures in complex formulas and visual grounding that existing printed-text benchmarks miss.

Motivation

Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.