Benchmark Radar
AI BENCHMARK PROFILE

FinRiskAtlas

Finance & EconomicsKnowledge & Reasoning

Evaluates Chinese financial LLMs on static operation execution across 9,742 instances and evidence-state control via FinRisk-Ask using pre-action trajectory states.

Released
2026-08-26
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Addresses the gap between broad financial benchmarks and workflow-specific decisions, showing that knowledge scores do not predict operational reliability.

Motivation

Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.