Benchmark Radar
AI BENCHMARK PROFILE

ITPEval

General AIMathematics & Formal SciencesITPEval team

ITPEval evaluates automated formal proof translation across four interactive theorem provers (Lean 4, Rocq, Isabelle, HOL Light), with 1,560 source files and 6,848 theorems, covering statement and proof translation on 12 directed pairs.

Released
2026-07-07
Readiness
Paper only
Primary field
General AI

Why it matters

ITPEval addresses the lack of a unified benchmark for cross-prover formal proof translation, providing a standardized evaluation of a key capability for automated reasoning and verified software portability.

Motivation

Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed across incompatible systems, limiting both training data for learning-based provers and the portability of verified results.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.