Benchmark Radar
AI BENCHMARK PROFILE

WhatIfBench

General AIKnowledge & Reasoning

A diagnostic benchmark of 220 open-domain what-if questions across STEM, HSS, and Hybrid scenarios, evaluated with PRISM metrics on causal graphs and explanatory adequacy.

Released
2026-08-28
Readiness
Runnable
Primary field
General AI

Why it matters

Exposes the gap between fluent counterfactual narrative and sound causal process reasoning, enabling more rigorous assessment of complex reasoning in LLMs.

Motivation

Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.