CFAgentBench
CFAgentBench is an executable environment for autonomous construction-finance agents, with 1,014 task specifications across 8 domains and 77 families. A subset of 40 tasks (54 with PM extension) has oracle-validated evaluators. Grading uses state diffs, forbidden-side-effect checks, and required-output regexes, with an LLM judge only for reply quality. A public split of 711 tasks is available, and a private split of 303 is reserved for remote scoring.
- Released
- 2026-06-20
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
The benchmark addresses the gap in evaluating agents for finance workflows that involve real software stacks and high-stakes transactions. Its focus on functional correctness and a money-movement guard, where correct actions can fail tasks, provides a more realistic measure of deployable competence. The observed performance collapse under repeated attempts highlights the need for reliability assessment beyond single-attempt accuracy.
Motivation
We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class agent operating across the real software stack a US construction finance team runs - ERP, project management, email, documents, pay applications, payroll, certified payroll, lien waivers, and bank/treasury portals.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.