Benchmark Radar
AI BENCHMARK PROFILE

ForecastBench-Sim

General AIKnowledge & Reasoning

A simulated-world forecasting benchmark built on Freeciv game rollouts. Evaluates probabilistic reasoning by scoring forecasts about hidden future game states, with continuous or binary questions, paired intervention worlds, and artifacts for release.

Released
2026-06-17
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the evaluation gap of slow real-world resolution and rare tail events by providing controllable, immediately resolvable forecasting tasks, enabling rigorous study of AI probabilistic reasoning under dynamic states.

Motivation

Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions are difficult to score.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.