Benchmark Radar
AI BENCHMARK PROFILE

OEIS Open

General AIKnowledge & ReasoningEpoch AI

Evaluates language models on formalized open mathematical conjectures from OEIS, scored by theorem resolution under a compute budget.

Released
2026-08-12
Readiness
Runnable
Primary field
General AI

Why it matters

It provides a secure, reproducible suite for testing autonomous mathematical reasoning on open problems.

Motivation

We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.