ClassicLogic
ClassicLogic is a benchmark suite of four classic logic puzzles (Sudoku, KenKen, Kakuro, Futoshiki) with a hierarchical knowledge base that defines complex strategies as compositions of simpler ones. It evaluates an agent's compositional generalization in problem-solving from basic rules to multi-step strategies across increasing puzzle difficulty.
- Released
- 2026-07-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Most compositional generalization benchmarks focus on language; ClassicLogic provides a structured, non-linguistic testbed with explicit compositional logic. It enables fine-grained assessment of reasoning capabilities and supports development of neuro-symbolic systems capable of systematic problem-solving.
Motivation
Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.