Benchmark Radar
AI BENCHMARK PROFILE

AnnoBench

General AIKnowledge & Reasoning

AnnoBench evaluates visualization annotation generation across four representation formats, five chart description conditions, and two prompt specification levels. It uses a VLM-as-a-judge protocol aligned with human assessment to score annotation quality.

Released
2026-07-28
Readiness
Paper only
Primary field
General AI

Why it matters

No existing benchmark tests whether annotation tools meet visual, semantic, and stylistic constraints. AnnoBench provides a structured evaluation framework to advance annotation automation and visualization generation pipelines.

Motivation

Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.