ComBench
ComBench evaluates large language models on 100 human-annotated Olympiad-level combinatorics problems, split into 50 analysis-centric and 50 construction-centric tasks. Scoring combines rubric-guided proof grading with deterministic verification of construction outputs.
- Released
- 2026-06-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
ComBench fills a gap in evaluating creative and rigorous combinatorial reasoning at the Olympiad level. It provides separate scores for proof quality and construction validity, helping diagnose where models diverge in these capabilities, and enables fine-grained comparison of frontier models.
Motivation
Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.