Benchmark Radar
AI BENCHMARK PROFILE

SDGBiasBench

General AIMultimodal Perception

SDGBiasBench is a benchmark suite for evaluating vision-language models on Sustainable Development Goals (SDG) reasoning, covering 500k multiple-choice questions and 50k regression tasks.

Released
2026-05-21
Readiness
Paper only
Primary field
General AI

Why it matters

Existing SDG evaluation tools lack a combined assessment of qualitative and quantitative reasoning, and this benchmark aims to expose systematic biases in model predictions. The value lies in enabling more reliable AI for sustainable development monitoring.

Motivation

Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.