Benchmark Radar
AI BENCHMARK PROFILE

TRACE-Bench

General AIMultimodal Perception

Evaluates multi-reference image generation by decomposing prompts into atomic operators (Anchor, Disentangle, Apply, Compose) and scoring operator-aligned capabilities and diagnostic failure localization.

Released
2026-08-17
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a capability-oriented protocol that captures combinatorial complexity and enables per-operator diagnostics, addressing the fragmented coverage and limited diagnostic value of task-type-based benchmarks.

Motivation

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.