Benchmark Radar
AI BENCHMARK PROFILE

MathAdv

General AIMathematics & Formal Sciences

Evaluates theorem proving and auxiliary tasks including multiple-choice, fill-in-the-blank, and reformulation robustness across 13 mathematical domains.

Released
2026-08-26
Readiness
Runnable
Primary field
General AI

Why it matters

Offers component-wise diagnosis of formal reasoning beyond aggregate proof accuracy, exposing formalization and robustness gaps.

Motivation

Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.