AI BENCHMARK PROFILE
MA-ProofBench
MA-ProofBench evaluates LLMs on theorem proving in mathematical analysis using 200 Lean 4 formalized problems split into undergraduate and Ph.D. levels. It covers six core topics and uses Pass@8 with formal verification via Lean server.
- Released
- 2026-06-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Formal theorem proving benchmarks typically lack coverage of advanced analysis. MA-ProofBench fills this gap, providing a stable evaluation for tracking progress in this difficult domain.
Motivation
Large Language Models (LLMs) have made notable progress in automated theorem proving, yet existing formal benchmarks remain limited in both mathematical coverage and difficulty.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.