Benchmark Radar
AI BENCHMARK PROFILE

FinPerMA

General AIKnowledge & Reasoning

FinPerMA evaluates personalized memory in LLM agents using frozen longitudinal investor trajectories with theory-informed impact rules. It includes a Post-Shock checkpoint to assess integration of material events into persistent user models, with 2,994 questions from 276 personas.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Existing personalized-memory benchmarks lack event-driven preference adaptation. FinPerMA fills this gap by providing an event-grounded evaluation, enabling assessment of whether agents can update user models over long horizons in high-stakes financial advising contexts.

Motivation

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.