Benchmark Radar
AI BENCHMARK PROFILE

PortBench

Finance & EconomicsKnowledge & ReasoningAgenticFinLab

PortBench evaluates LLM-driven portfolio management via a static QA dataset (6,269 questions across seven task templates) and a dynamic five-stage allocation pipeline, spanning six asset classes over ten years. Scoring includes a dual-layer correlation score and CEPS, with evaluation under three stress regimes and investor profiles.

Released
2026-05-27
Readiness
Runnable
Primary field
Finance & Economics

Why it matters

Existing financial benchmarks often ignore cross-asset correlations and the full portfolio management pipeline. PortBench addresses this gap by providing a comprehensive, reusable evaluation for LLM capabilities in realistic portfolio management, enabling comparison across models and informing practical deployment decisions.

Motivation

Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.