Benchmark Radar
AI BENCHMARK PROFILE

TheoremBench

General AIMathematics & Formal SciencesTheoremBench Team

Evaluates LLMs on theorem proving in Lean4 using classical theorems, with two versions: main and premised. Includes metrics for theorem-level coverage and token efficiency to assess partial progress and proof structure.

Released
2026-06-08
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a more realistic evaluation of provers beyond contest problems, revealing biases toward easy subtheorems and inefficient proof strategies. Supports finer-grained analysis of formal reasoning capabilities.

Motivation

LLMs have recently achieved strong results on formal proving benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.