Benchmark Radar
AI BENCHMARK PROFILE

CombEval

General AIKnowledge & ReasoningYuxuZhou-CN

CombEval evaluates combinatorial counting abilities of large language models using problems generated from typed Cofola specifications, with solver-verified answers. It supports systematic variation of object types, entity scales, constraint counts, and reasoning depth.

Released
2026-06-18
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap in dynamic evaluation of combinatorial reasoning, providing controlled generation and exact verification to diagnose model failures in counting tasks, useful for targeted improvements.

Motivation

We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.