AI BENCHMARK PROFILE
PyraMathBench
PyraMathBench is a hierarchical benchmark with 32,505 questions derived from math word problems, evaluating numerical reasoning and mathematical capabilities across cognitive aspects and modalities.
- Released
- 2026-06-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
No public artifacts or reuse path are provided, and the benchmark lacks a clear scoring contract for independent runs.
Motivation
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.