Benchmark Radar
AI BENCHMARK PROFILE

CausalGame

General AIAgentsCausalGame Team

CausalGame is a benchmark for evaluating causal thinking in LLM agents through interactive games. It includes 14 scenarios with selection bias, measurement error, and hidden confounders. Agents design experimental protocols, collect data, and produce solutions with explanations, scored against analytical optima and causal-reasoning rubrics.

Released
2026-07-05
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing AI Scientist benchmarks do not isolate causal reasoning under realistic biases. CausalGame provides a structured evaluation for this capability, offering decision value for developers of autonomous research agents seeking to assess robustness against confounders and biases.

Motivation

Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.