Benchmark Radar
AI BENCHMARK PROFILE

SalArt-VQA

General AIMultimodal PerceptionSalArt-VQA Benchmark Team

SalArt-VQA is a closed-set visual question answering benchmark for fine-grained salient artifact understanding in AI-generated images. It includes 950 images and 3,681 human-authored multiple-choice questions covering artifact images, matched real references, and paired generated references. Four aligned question types evaluate presence detection, semantic localization, spatial grounding, and evidence-grounded defect identification, with reference splits for calibration and abstention testing.

Released
2026-06-10
Readiness
Inspectable
Primary field
General AI

Why it matters

Image-level artifact detection accuracy can conceal failures in grounding and evidence use. This benchmark provides a fine-grained evaluation protocol that isolates specific failure modes, enabling comparison of VLMs on their ability to support artifact claims with local visual evidence. It offers practical value for developers selecting models for trustworthy artifact analysis in generated image workflows.

Motivation

Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.