Benchmark Radar
AI BENCHMARK PROFILE

FinPersona-Bench

Finance & EconomicsAgents

Evaluates the longitudinal stability of behavioral mandates in LLM-based financial agents using a synthetic market simulation that decouples observable price from hidden fundamental value, scoring mandate adherence across calm, crash, and bubble market regimes.

Released
2026-06-30
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

It addresses the gap in evaluating long-horizon behavioral consistency of autonomous financial agents, providing a falsifiable method to measure mandate salience decay and informing deployment decisions for mandate-aware re-grounding strategies.

Motivation

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.