Benchmark Radar
AI BENCHMARK PROFILE

CraftBench

General AIMultimodal PerceptionHaozheZhao/Crafter

CraftBench evaluates scientific figure generation across three figure types and four input conditions, with human-drawn targets and a referenced VLM judge for scoring.

Released
2026-05-28
Readiness
Runnable
Primary field
General AI

Why it matters

Automated figure generation lacks comprehensive benchmarks covering diverse types and conditions. CraftBench provides a standardized evaluation to measure progress in editable scientific figure generation.

Motivation

Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensive parts of paper preparation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.