Benchmark Radar
AI BENCHMARK PROFILE

AtelierEval

General AIMultimodal Perception

AtelierEval evaluates prompting proficiency of humans and MLLMs for text-to-image systems via 360 tasks, using a skill-based agentic evaluator (AtelierJudge) that scores prompt-image pairs subjectively and objectively.

Released
2026-05-21
Readiness
Paper only
Primary field
General AI

Why it matters

Introduces a new evaluation angle for T2I pipelines, measuring upstream prompting ability which is currently unassessed.

Motivation

Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.