ChLogic
ChLogic is an English–Chinese aligned benchmark for evaluating logical reasoning robustness across surface realizations. It includes 3,000 general, 2,000 difficult, and 1,500 Chinese-only items derived from formal logical templates, with each aligned item pairing one English reference with five Chinese variants.
- Released
- 2026-06-16
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing logical reasoning benchmarks focus on English, so it is unclear whether models retain performance when the same logical structure is expressed in Chinese. ChLogic provides a cross-lingual stress test, helping identify language-specific gaps and translation artifacts that affect multilingual reasoning.
Motivation
Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.