Benchmark Radar
AI BENCHMARK PROFILE

Edit2TikZ

General AIMultimodal PerceptionCoding & Software EngineeringSolunny

Edit2TikZ evaluates instruction-guided scientific figure editing with TikZ code, featuring 1,548 samples with textual or visual localization requests and multi-step edits, using a human-aligned evaluation framework to measure edit completion and content preservation.

Released
2026-08-13
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on figure reconstruction or generation, leaving a gap for systematic evaluation of instruction-guided editing with compilable code. This benchmark provides a standardized protocol for assessing models' ability to perform precise, code-based edits while preserving unrelated content, offering practical value for developing reliable multimodal systems.

Motivation

Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground the requested change, generate compilable code, and preserve all unrelated content.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.