Principle-Bench
Principle-Bench contains 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations for evaluating LLM-as-judge on accuracy, paraphrase robustness, adversarial robustness, and calibration.
- Released
- 2026-08-14
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Addresses the evaluation gap for LLM-as-judge in principle-based regulation, where standards are not binary. Provides a multi-axis assessment to inform deployment decisions, as no single method dominates all axes and adversarial inputs can significantly degrade performance.
Motivation
Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.