Benchmark Radar
AI BENCHMARK PROFILE

InvestLogicBench

Finance & EconomicsKnowledge & Reasoning

InvestLogicBench evaluates large language models on personalized investment decision-making using 201,247 documented decisions from 151 real-world investors. Each episode traces investor profile, market events, reasoning, decision, and outcome. Tasks include comprehension, profile-conditioned generation, and end-to-end replay.

Released
2026-08-06
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Existing financial LLM evaluations rely on static QA or terminal profit, which fail to reveal whether actions are profile-consistent or grounded in events. This benchmark targets a gap by assessing process quality and grounding in personalized, consequential settings.

Motivation

Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.