BigFinanceBench
BigFinanceBench evaluates financial-research agents on open-ended tasks with 928 items, each paired with a ground-truth answer and a point-weighted rubric decomposing the derivation into steps. Supports partial-credit scoring across 36,241 rubric points.
- Released
- 2026-06-02
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Existing finance benchmarks evaluate subskills or final answers, not the auditable derivation. BigFinanceBench measures workflow quality, allowing localization of failures and better assessment of decision-relevant outputs.
Motivation
Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.