Benchmark Radar
AI BENCHMARK PROFILE

CFAgentBench

Finance & EconomicsAgentsCFAgentBench Team

CFAgentBench is an executable environment for autonomous construction-finance agents, with 1,014 task specifications across 8 domains and 77 families. A subset of 40 tasks (54 with PM extension) has oracle-validated evaluators. Grading uses state diffs, forbidden-side-effect checks, and required-output regexes, with an LLM judge only for reply quality. A public split of 711 tasks is available, and a private split of 303 is reserved for remote scoring.

Released
2026-06-20
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

The benchmark addresses the gap in evaluating agents for finance workflows that involve real software stacks and high-stakes transactions. Its focus on functional correctness and a money-movement guard, where correct actions can fail tasks, provides a more realistic measure of deployable competence. The observed performance collapse under repeated attempts highlights the need for reliability assessment beyond single-attempt accuracy.

Motivation

We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class agent operating across the real software stack a US construction finance team runs - ERP, project management, email, documents, pay applications, payroll, certified payroll, lien waivers, and bank/treasury portals.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.