AI BENCHMARK PROFILE
CombEval
CombEval evaluates combinatorial counting abilities of large language models using problems generated from typed Cofola specifications, with solver-verified answers. It supports systematic variation of object types, entity scales, constraint counts, and reasoning depth.
- Released
- 2026-06-18
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap in dynamic evaluation of combinatorial reasoning, providing controlled generation and exact verification to diagnose model failures in counting tasks, useful for targeted improvements.
Motivation
We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.