Benchmark Radar
AI BENCHMARK PROFILE

ClassicLogic

General AIKnowledge & ReasoningClassicLogic Team

ClassicLogic is a benchmark suite of four classic logic puzzles (Sudoku, KenKen, Kakuro, Futoshiki) with a hierarchical knowledge base that defines complex strategies as compositions of simpler ones. It evaluates an agent's compositional generalization in problem-solving from basic rules to multi-step strategies across increasing puzzle difficulty.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

Most compositional generalization benchmarks focus on language; ClassicLogic provides a structured, non-linguistic testbed with explicit compositional logic. It enables fine-grained assessment of reasoning capabilities and supports development of neuro-symbolic systems capable of systematic problem-solving.

Motivation

Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.