Benchmark Radar
AI BENCHMARK PROFILE

InvestPhilBench

Finance & EconomicsKnowledge & Reasoning

InvestPhilBench is a multi-layer benchmark for evaluating procedural reasoning in investment philosophy, spanning eight cognitive tiers from principle identification to novel framework extrapolation. It includes principle cards, decision-framework cards, and QA questions, with an automated scoring pipeline (BASP) and five algorithmic metrics.

Released
2026-06-24
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Fills the gap in testing whether LLMs can accurately reconstruct and apply expert procedural decision frameworks. Provides a reproducible method for scoring procedural reasoning, with a metric (GRA) that exposes deficits hidden by composite scores.

Motivation

Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.