Benchmark Radar
AI BENCHMARK PROFILE

CPI-Bench

General AIMultimodal PerceptionTaobaoTmall-AlgorithmProducts

CPI-Bench evaluates image editing models across general, practical, and intelligent tasks with VLM-as-Judge scoring, including multi-image and reasoning-based editing.

Released
2026-08-14
Readiness
Runnable
Primary field
General AI

Why it matters

It provides a comprehensive benchmark that captures real-world deployment scenarios and reasoning demands, with scoring aligned to human preferences for reliable model comparison.

Motivation

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.