AnnoBench
AnnoBench evaluates visualization annotation generation across four representation formats, five chart description conditions, and two prompt specification levels. It uses a VLM-as-a-judge protocol aligned with human assessment to score annotation quality.
- Released
- 2026-07-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
No existing benchmark tests whether annotation tools meet visual, semantic, and stylistic constraints. AnnoBench provides a structured evaluation framework to advance annotation automation and visualization generation pipelines.
Motivation
Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.