PortBench
PortBench evaluates LLM-driven portfolio management via a static QA dataset (6,269 questions across seven task templates) and a dynamic five-stage allocation pipeline, spanning six asset classes over ten years. Scoring includes a dual-layer correlation score and CEPS, with evaluation under three stress regimes and investor profiles.
- Released
- 2026-05-27
- Readiness
- Runnable
- Primary field
- Finance & Economics
Why it matters
Existing financial benchmarks often ignore cross-asset correlations and the full portfolio management pipeline. PortBench addresses this gap by providing a comprehensive, reusable evaluation for LLM capabilities in realistic portfolio management, enabling comparison across models and informing practical deployment decisions.
Motivation
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.