InvestLogicBench
InvestLogicBench evaluates large language models on personalized investment decision-making using 201,247 documented decisions from 151 real-world investors. Each episode traces investor profile, market events, reasoning, decision, and outcome. Tasks include comprehension, profile-conditioned generation, and end-to-end replay.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Existing financial LLM evaluations rely on static QA or terminal profit, which fail to reveal whether actions are profile-consistent or grounded in events. This benchmark targets a gap by assessing process quality and grounding in personalized, consequential settings.
Motivation
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.