Benchmark Radar
AI BENCHMARK PROFILE

SDABench

Science & ResearchKnowledge & Reasoning

SDABench evaluates LLMs' scientific data analysis capabilities across six capabilities (descriptive, exploratory, inferential, predictive, causal, mechanistic) and five domains, with 527 real and 6000 synthetic instances in multiple-choice and open-ended formats.

Released
2026-07-13
Readiness
Paper only
Primary field
Science & Research

Why it matters

This benchmark reveals that LLMs degrade sharply on tasks requiring assumption selection, latent-process modeling, and mechanistic reasoning, highlighting gaps for scientific discovery applications.

Motivation

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.