Benchmark Radar
AI BENCHMARK PROFILE

FinBench

Finance & EconomicsSafety & Trustworthiness

FinBench is a benchmark for evaluating calibration and uncertainty in financial forecasting with time-gated tasks, requiring probability of positive return and 80% prediction interval, scored with Brier and Winkler scores.

Released
2026-06-24
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Financial forecasting agents risk overconfidence; FinBench addresses this by testing probabilistic calibration under temporal constraints, but its pilot scale limits immediate comparison value.

Motivation

Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.