Benchmark Radar
AI BENCHMARK PROFILE

SciDraw-Bench

General AIMultimodal Perception

SciDraw-Bench is a benchmark for scientific figure generation, with 32 tasks across eight figure types and ten disciplines. Each task pairs a natural-language prompt with a machine-checkable specification, and evaluation uses four dimensions: Text Fidelity, Semantic Correctness, Structural Quality, and Convention Adherence.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Fills the gap in evaluating scientific figure generation, which requires correct labels and diagrammatic structure beyond natural-image composition. It provides a protocol for measuring usability of generated figures, aiding development of domain-specific generative systems.

Motivation

Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.