Benchmark Radar
AI BENCHMARK PROFILE

PyraMathBench

General AIMathematics & Formal Sciences

PyraMathBench is a hierarchical benchmark with 32,505 questions derived from math word problems, evaluating numerical reasoning and mathematical capabilities across cognitive aspects and modalities.

Released
2026-06-02
Readiness
Paper only
Primary field
General AI

Why it matters

No public artifacts or reuse path are provided, and the benchmark lacks a clear scoring contract for independent runs.

Motivation

Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.