Benchmark Radar
AI BENCHMARK PROFILE

SciFigQual-Bench

General AIMultimodal PerceptionFrankDengAI

A benchmark for evaluating scientific figure quality across five dimensions with full-manuscript context, using 6,308 expert-rated images from top CS conferences and a fixed eval1200 test split.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

Existing image quality assessment methods are unsuitable for scientific figures; this benchmark provides a contextual, multi-dimensional standard for automated evaluation, enabling model comparison.

Motivation

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.