Benchmark Radar
AI BENCHMARK PROFILE

FORCE-Bench

Finance & EconomicsKnowledge & ReasoningFORCE-Bench team

FORCE-Bench evaluates agentic AI systems in enterprise finance across three task types: financial obligation research, financial entity performance research, and business brief generation. It includes 251 expert-annotated queries and a rubric-based scoring framework across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure.

Released
2026-07-11
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Existing benchmarks focus on general capabilities rather than operational finance workflows. FORCE-Bench provides a domain-specific evaluation tool that measures rule adherence, verifiability, and groundedness, which are critical for real-world deployment of agentic systems in finance. It enables comparable assessment of agent performance under operational constraints.

Motivation

Recent advances in large language models have accelerated deployment of agentic systems in operational finance.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.