Benchmark Radar
AI BENCHMARK PROFILE

BlueFin

Robotics & Autonomous SystemsFinance & EconomicsRobotics & Embodied IntelligenceBlueFin Benchmark Team

Evaluates LLM agents on synthesis, manipulation, and comprehension tasks over financial spreadsheet workbooks, with 131 tasks and granular rubric criteria validated by expert annotators.

Released
2026-05-29
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Fills the gap in evaluating LLMs for spreadsheet tasks relevant to professional finance, where current models perform below 50%, providing a measure for practical deployment.

Motivation

We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet workbooks in the professional finance domain.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.