AI BENCHMARK PROFILE
MathAdv
Evaluates theorem proving and auxiliary tasks including multiple-choice, fill-in-the-blank, and reformulation robustness across 13 mathematical domains.
- Released
- 2026-08-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Offers component-wise diagnosis of formal reasoning beyond aggregate proof accuracy, exposing formalization and robustness gaps.
Motivation
Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.