SDABench
SDABench evaluates LLMs' scientific data analysis capabilities across six capabilities (descriptive, exploratory, inferential, predictive, causal, mechanistic) and five domains, with 527 real and 6000 synthetic instances in multiple-choice and open-ended formats.
- Released
- 2026-07-13
- Readiness
- Paper only
- Primary field
- Science & Research
Why it matters
This benchmark reveals that LLMs degrade sharply on tasks requiring assumption selection, latent-process modeling, and mechanistic reasoning, highlighting gaps for scientific discovery applications.
Motivation
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.