Benchmark Radar
AI BENCHMARK PROFILE

LL-Bench

General AIMultimodal Perception

LL-Bench evaluates large-scale generative models on 16 low-level vision tasks using 2,469 real-world degraded images, with human preference and quality score annotations for model outputs.

Released
2026-06-01
Readiness
Paper only
Primary field
General AI

Why it matters

Existing low-level vision benchmarks often focus on conventional models and lack alignment with human perception. LL-Bench provides a standardized evaluation suite to compare generative and restoration models on pixel-level tasks, supporting quality assessment and model selection.

Motivation

Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.