Benchmark Radar
AI BENCHMARK PROFILE

TangPoetryBench

General AIMultimodal Perception

TangPoetryBench evaluates text-to-image models on illustrating classical Chinese Tang poems across ten human-annotated dimensions, with a rubric-conditioned evaluator (PAE).

Released
2026-08-11
Readiness
Paper only
Primary field
General AI

Why it matters

Existing metrics fail to capture cultural and emotional fidelity in poetry-to-image generation; this benchmark provides a multi-dimensional human-annotated dataset and an automated evaluator to support model comparison in this niche domain.

Motivation

Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.