MPIE-Bench
MPIE-Bench evaluates multi-person image editing models on tasks involving multiple named people in contact interactions such as embrace, carry, or grapple. The benchmark provides a 2,500-sample test set with 14 interaction categories and four contact densities, and scores outputs on six axes including anatomy and interaction via mesh reconstruction.
- Released
- 2026-07-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing evaluations often overlook anatomical and geometric errors in multi-person editing, and VLM-based judges may rate such errors as acceptable. This benchmark provides a geometry-based scoring method that tracks human judgment more closely, enabling more reliable comparisons of editing models on this challenging task.
Motivation
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.