Benchmark Radar
AI BENCHMARK PROFILE

Fiducia-bench

Finance & EconomicsKnowledge & Reasoning

Evaluates governability of financial agents across 626 episodes and 100 KYC/AML task variants, measuring escalation, abstention, and audit trail completeness under different agent architectures.

Released
2026-08-17
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Identifies a governance gap in decomposed agent architectures, providing a framework to assess whether policy compliance is preserved when agents are broken into components.

Motivation

Existing agent benchmarks ask whether the agent finished the task.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.