SalArt-VQA
SalArt-VQA is a closed-set visual question answering benchmark for fine-grained salient artifact understanding in AI-generated images. It includes 950 images and 3,681 human-authored multiple-choice questions covering artifact images, matched real references, and paired generated references. Four aligned question types evaluate presence detection, semantic localization, spatial grounding, and evidence-grounded defect identification, with reference splits for calibration and abstention testing.
- Released
- 2026-06-10
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Image-level artifact detection accuracy can conceal failures in grounding and evidence use. This benchmark provides a fine-grained evaluation protocol that isolates specific failure modes, enabling comparison of VLMs on their ability to support artifact claims with local visual evidence. It offers practical value for developers selecting models for trustworthy artifact analysis in generated image workflows.
Motivation
Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.