AI BENCHMARK PROFILE
Diagram-MMU
Diagram-MMU evaluates multimodal language models on parsing scientific diagrams into LaTeX TikZ code, editing diagram code, and answering questions about diagrams, using 3.7k diagrams and 18.3k questions across six domains.
- Released
- 2026-08-12
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
The evaluation provides insight into model capabilities on diagram-to-code generation tasks, a practical need in scientific authoring tools, and identifies gaps between reasoning and code generation.
Motivation
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.