Benchmark Radar
AI BENCHMARK PROFILE

ChLogic

General AISafety & TrustworthinessChLogic team

ChLogic is an English–Chinese aligned benchmark for evaluating logical reasoning robustness across surface realizations. It includes 3,000 general, 2,000 difficult, and 1,500 Chinese-only items derived from formal logical templates, with each aligned item pairing one English reference with five Chinese variants.

Released
2026-06-16
Readiness
Runnable
Primary field
General AI

Why it matters

Existing logical reasoning benchmarks focus on English, so it is unclear whether models retain performance when the same logical structure is expressed in Chinese. ChLogic provides a cross-lingual stress test, helping identify language-specific gaps and translation artifacts that affect multilingual reasoning.

Motivation

Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.