AI BENCHMARK PROFILE
FinRiskAtlas
Evaluates Chinese financial LLMs on static operation execution across 9,742 instances and evidence-state control via FinRisk-Ask using pre-action trajectory states.
- Released
- 2026-08-26
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Addresses the gap between broad financial benchmarks and workflow-specific decisions, showing that knowledge scores do not predict operational reliability.
Motivation
Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.