ForecastBench-Sim
A simulated-world forecasting benchmark built on Freeciv game rollouts. Evaluates probabilistic reasoning by scoring forecasts about hidden future game states, with continuous or binary questions, paired intervention worlds, and artifacts for release.
- Released
- 2026-06-17
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the evaluation gap of slow real-world resolution and rare tail events by providing controllable, immediately resolvable forecasting tasks, enabling rigorous study of AI probabilistic reasoning under dynamic states.
Motivation
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions are difficult to score.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.