Benchmark Radar
AI BENCHMARK PROFILE

ComBench

General AIMathematics & Formal SciencesSimplified Reasoning

ComBench evaluates large language models on 100 human-annotated Olympiad-level combinatorics problems, split into 50 analysis-centric and 50 construction-centric tasks. Scoring combines rubric-guided proof grading with deterministic verification of construction outputs.

Released
2026-06-09
Readiness
Runnable
Primary field
General AI

Why it matters

ComBench fills a gap in evaluating creative and rigorous combinatorial reasoning at the Olympiad level. It provides separate scores for proof quality and construction validity, helping diagnose where models diverge in these capabilities, and enables fine-grained comparison of frontier models.

Motivation

Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.