FORCE-Bench
FORCE-Bench evaluates agentic AI systems in enterprise finance across three task types: financial obligation research, financial entity performance research, and business brief generation. It includes 251 expert-annotated queries and a rubric-based scoring framework across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure.
- Released
- 2026-07-11
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Existing benchmarks focus on general capabilities rather than operational finance workflows. FORCE-Bench provides a domain-specific evaluation tool that measures rule adherence, verifiability, and groundedness, which are critical for real-world deployment of agentic systems in finance. It enables comparable assessment of agent performance under operational constraints.
Motivation
Recent advances in large language models have accelerated deployment of agentic systems in operational finance.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.