Benchmark Radar
AI BENCHMARK PROFILE

BigFinanceBench

Finance & EconomicsKnowledge & Reasoning

BigFinanceBench evaluates financial-research agents on open-ended tasks with 928 items, each paired with a ground-truth answer and a point-weighted rubric decomposing the derivation into steps. Supports partial-credit scoring across 36,241 rubric points.

Released
2026-06-02
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Existing finance benchmarks evaluate subskills or final answers, not the auditable derivation. BigFinanceBench measures workflow quality, allowing localization of failures and better assessment of decision-relevant outputs.

Motivation

Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.