Benchmark Radar
AI BENCHMARK PROFILE

AdvancedMathBench

General AIMathematics & Formal Sciences

AdvancedMathBench is a benchmark suite for advanced mathematical reasoning, containing ProverBench (296 proof problems) and VerifierBench (888 proof trajectories with expert labels), with an automatic verification pipeline.

Released
2026-07-13
Readiness
Paper only
Primary field
General AI

Why it matters

Existing math benchmarks focus on high-school levels and final answers. AdvancedMathBench evaluates proof generation and verification at advanced levels, providing fine-grained assessments of proof correctness and error detection.

Motivation

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.