AI BENCHMARK PROFILE
TangPoetryBench
TangPoetryBench evaluates text-to-image models on illustrating classical Chinese Tang poems across ten human-annotated dimensions, with a rubric-conditioned evaluator (PAE).
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing metrics fail to capture cultural and emotional fidelity in poetry-to-image generation; this benchmark provides a multi-dimensional human-annotated dataset and an automated evaluator to support model comparison in this niche domain.
Motivation
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.