Benchmark Radar
AI BENCHMARK PROFILE

SpreadsheetBench

Finance & EconomicsCoding & Software EngineeringRUCKBReasoning

SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business workflows across generation, debugging, and visualization tasks. It includes 321 tasks from authentic business data, with multi-sheet workbooks requiring cross-sheet reasoning. The benchmark provides a unified multi-turn agent scaffold and evaluation scripts for reproducibility.

Released
2026-06-29
Readiness
Runnable
Primary field
Finance & Economics

Why it matters

Existing spreadsheet benchmarks focus on isolated operations, failing to capture real-world workflow complexity. SpreadsheetBench 2 addresses this gap by assessing agents on tasks that require multi-step coordination, cross-sheet reasoning, and deliverable-level outcomes. It provides a challenging testbed for improving reliable spreadsheet automation, with current models achieving only 34.89% overall accuracy.

Motivation

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.