Benchmark Radar
AI BENCHMARK PROFILE

MA-ProofBench

General AIMathematics & Formal SciencesOpenBMB

MA-ProofBench evaluates LLMs on theorem proving in mathematical analysis using 200 Lean 4 formalized problems split into undergraduate and Ph.D. levels. It covers six core topics and uses Pass@8 with formal verification via Lean server.

Released
2026-06-11
Readiness
Runnable
Primary field
General AI

Why it matters

Formal theorem proving benchmarks typically lack coverage of advanced analysis. MA-ProofBench fills this gap, providing a stable evaluation for tracking progress in this difficult domain.

Motivation

Large Language Models (LLMs) have made notable progress in automated theorem proving, yet existing formal benchmarks remain limited in both mathematical coverage and difficulty.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.