Benchmark Radar
AI BENCHMARK PROFILE

SABER-Math

General AIKnowledge & ReasoningSearch & RetrievalMathematics & Formal Sciences

SABER-Math evaluates information retrieval for mathematical queries, with about 283K problems and tasks for reranking based on fine-grained relevance. Scoring uses preference tournament ratings.

Released
2026-06-29
Readiness
Paper only
Primary field
General AI

Why it matters

Existing IR benchmarks fail to capture mathematical relevance, and MTEB doesn't predict math performance. SABER-Math provides a math-specific benchmark to guide retriever selection.

Motivation

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.