Dr-CiK
Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant supporting context from a document corpus, filter distractors, distill evidence, and generate forecasts. It provides context ablations and evaluates deep research and forecasting methods.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Real-world forecasting requires active discovery of external context from noisy sources. Dr-CiK assesses the entire pipeline of context retrieval and use, revealing that most agents recover little evidence and are misled by distractors, guiding development of foresight-driven agents.
Motivation
Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.