Benchmark Radar
AI BENCHMARK PROFILE

SocSci-Repro-Bench

General AIKnowledge & Reasoning

SocSci-Repro-Bench evaluates AI coding agents on reproducing social science findings from 221 tasks across four disciplines and 13 domains, using studies with known reproducibility outcomes.

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

SocSci-Repro-Bench addresses the lack of systematic evaluation of AI agents' ability to execute computational workflows, providing a protocol to isolate agent performance from issues in reproduction materials and informing practical use of agents in scientific production.

Motivation

Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provided with original data and code; yet systematic evaluation across social sciences remains limited.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.