Benchmark Radar
AI BENCHMARK PROFILE

PaintBench

Robotics & Autonomous SystemsRobotics & Embodied IntelligenceNYU

PaintBench is a procedurally generated benchmark for precise visual editing, covering 20 tasks across geometric, structural, color, and symbolic categories. Stable scoring uses pixel-level mIoU with fixed seeds.

Released
2026-05-29
Readiness
Runnable
Primary field
Robotics & Autonomous Systems

Why it matters

Procedural generation allows contamination-resistant evaluation of precise editing capabilities, with deterministic scoring that avoids human or LLM bias.

Motivation

While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.