SDGBiasBench
SDGBiasBench is a benchmark suite for evaluating vision-language models on Sustainable Development Goals (SDG) reasoning, covering 500k multiple-choice questions and 50k regression tasks.
- Released
- 2026-05-21
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing SDG evaluation tools lack a combined assessment of qualitative and quantitative reasoning, and this benchmark aims to expose systematic biases in model predictions. The value lies in enabling more reliable AI for sustainable development monitoring.
Motivation
Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.