Benchmark Radar
AI BENCHMARK PROFILE

MPIE-Bench

General AIMultimodal PerceptionMPIE-Bench Team

MPIE-Bench evaluates multi-person image editing models on tasks involving multiple named people in contact interactions such as embrace, carry, or grapple. The benchmark provides a 2,500-sample test set with 14 interaction categories and four contact densities, and scores outputs on six axes including anatomy and interaction via mesh reconstruction.

Released
2026-07-30
Readiness
Runnable
Primary field
General AI

Why it matters

Existing evaluations often overlook anatomical and geometric errors in multi-person editing, and VLM-based judges may rate such errors as acceptable. This benchmark provides a geometry-based scoring method that tracks human judgment more closely, enabling more reliable comparisons of editing models on this challenging task.

Motivation

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.