SocSci-Repro-Bench
SocSci-Repro-Bench evaluates AI coding agents on reproducing social science findings from 221 tasks across four disciplines and 13 domains, using studies with known reproducibility outcomes.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
SocSci-Repro-Bench addresses the lack of systematic evaluation of AI agents' ability to execute computational workflows, providing a protocol to isolate agent performance from issues in reproduction materials and informing practical use of agents in scientific production.
Motivation
Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provided with original data and code; yet systematic evaluation across social sciences remains limited.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.