CausalGame
CausalGame is a benchmark for evaluating causal thinking in LLM agents through interactive games. It includes 14 scenarios with selection bias, measurement error, and hidden confounders. Agents design experimental protocols, collect data, and produce solutions with explanations, scored against analytical optima and causal-reasoning rubrics.
- Released
- 2026-07-05
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing AI Scientist benchmarks do not isolate causal reasoning under realistic biases. CausalGame provides a structured evaluation for this capability, offering decision value for developers of autonomous research agents seeking to assess robustness against confounders and biases.
Motivation
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.