AI BENCHMARK PROFILE
BlueFin
Evaluates LLM agents on synthesis, manipulation, and comprehension tasks over financial spreadsheet workbooks, with 131 tasks and granular rubric criteria validated by expert annotators.
- Released
- 2026-05-29
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Fills the gap in evaluating LLMs for spreadsheet tasks relevant to professional finance, where current models perform below 50%, providing a measure for practical deployment.
Motivation
We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet workbooks in the professional finance domain.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.