AI BENCHMARK PROFILE
LL-Bench
LL-Bench evaluates large-scale generative models on 16 low-level vision tasks using 2,469 real-world degraded images, with human preference and quality score annotations for model outputs.
- Released
- 2026-06-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing low-level vision benchmarks often focus on conventional models and lack alignment with human perception. LL-Bench provides a standardized evaluation suite to compare generative and restoration models on pixel-level tasks, supporting quality assessment and model selection.
Motivation
Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.