FinGuard-Bench
The evaluation focuses on financial regulatory non-compliance detection in LLM interactions. It involves a benchmark with expert-annotated labels at query and response levels, but the exact task, environment, or scoring setup is not detailed.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Assesses LLM compliance with financial regulations, which is critical for preventing regulatory penalties and consumer harm in financial services. The benchmark aims to measure detection capabilities across institution-specific policies.
Motivation
As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulatory penalties and direct consumer harm.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.