Benchmark Radar
AI BENCHMARK PROFILE

CausalDS

General AIKnowledge & Reasoning

Evaluates causal reasoning in data-science workflows using synthetic scenes with hidden structural causal models, covering Pearl's three rungs and abstention scoring.

Released
2026-07-09
Readiness
Runnable
Primary field
General AI

Why it matters

Bridges symbolic causal reasoning and realistic data analysis, providing a joint evaluation of reasoning, tool use, and uncertainty quantification in agentic settings.

Motivation

Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.