Benchmark Radar
AI BENCHMARK PROFILE

Diagram-MMU

General AIMultimodal Perception

Diagram-MMU evaluates multimodal language models on parsing scientific diagrams into LaTeX TikZ code, editing diagram code, and answering questions about diagrams, using 3.7k diagrams and 18.3k questions across six domains.

Released
2026-08-12
Readiness
Inspectable
Primary field
General AI

Why it matters

The evaluation provides insight into model capabilities on diagram-to-code generation tasks, a practical need in scientific authoring tools, and identifies gaps between reasoning and code generation.

Motivation

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.