AI BENCHMARK PROFILE
FaithformBench
A benchmark for evaluating faithfulness of mathematical chain-of-thought autoformalisation, using perturbed reasoning steps to test validity and invalidity preservation.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the need for sound and cheap evaluation of autoformalisation faithfulness, revealing sycophancy in current systems.
Motivation
Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.