FinPersona-Bench
Evaluates the longitudinal stability of behavioral mandates in LLM-based financial agents using a synthetic market simulation that decouples observable price from hidden fundamental value, scoring mandate adherence across calm, crash, and bubble market regimes.
- Released
- 2026-06-30
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
It addresses the gap in evaluating long-horizon behavioral consistency of autonomous financial agents, providing a falsifiable method to measure mandate salience decay and informing deployment decisions for mandate-aware re-grounding strategies.
Motivation
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.