AI BENCHMARK PROFILE
WhatIfBench
A diagnostic benchmark of 220 open-domain what-if questions across STEM, HSS, and Hybrid scenarios, evaluated with PRISM metrics on causal graphs and explanatory adequacy.
- Released
- 2026-08-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Exposes the gap between fluent counterfactual narrative and sound causal process reasoning, enabling more rigorous assessment of complex reasoning in LLMs.
Motivation
Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.