Benchmark Radar
AI BENCHMARK PROFILE

FM-Bench

Finance & EconomicsKnowledge & ReasoningAnalogy AI

FM-Bench evaluates LLM agents on long-horizon football club management through 20 in-game years with 26 tools, measuring managerial decision quality via a deterministic engine scoring.

Released
2026-08-19
Readiness
Runnable
Primary field
Finance & Economics

Why it matters

Current agent benchmarks focus on bounded tasks, leaving long-horizon decision-making unmeasured. FM-Bench provides a reproducible environment with deterministic scoring to test sustained decision quality over hundreds of steps.

Motivation

Language model agents now execute bounded tasks reliably.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.