Benchmark Radar
AI BENCHMARK PROFILE

FaithformBench

General AIKnowledge & Reasoning

A benchmark for evaluating faithfulness of mathematical chain-of-thought autoformalisation, using perturbed reasoning steps to test validity and invalidity preservation.

Released
2026-08-11
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the need for sound and cheap evaluation of autoformalisation faithfulness, revealing sycophancy in current systems.

Motivation

Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.