AI BENCHMARK PROFILE
CausalDS
Evaluates causal reasoning in data-science workflows using synthetic scenes with hidden structural causal models, covering Pearl's three rungs and abstention scoring.
- Released
- 2026-07-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Bridges symbolic causal reasoning and realistic data analysis, providing a joint evaluation of reasoning, tool use, and uncertainty quantification in agentic settings.
Motivation
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.