AI BENCHMARK PROFILE
FinBench
FinBench is a benchmark for evaluating calibration and uncertainty in financial forecasting with time-gated tasks, requiring probability of positive return and 80% prediction interval, scored with Brier and Winkler scores.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Financial forecasting agents risk overconfidence; FinBench addresses this by testing probabilistic calibration under temporal constraints, but its pilot scale limits immediate comparison value.
Motivation
Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.