Benchmark Radar
AI BENCHMARK PROFILE

ChaosBench-Logic

General AIKnowledge & Reasoning

ChaosBench-Logic v2 evaluates LLM logical reasoning over dynamical systems with 40,886 questions across 165 systems, 27 first-order logic predicates, and 78 axiom edges. The CARE protocol measures calibration and adversarial robustness, reporting metrics such as MCC.

Released
2026-05-23
Readiness
Paper only
Primary field
General AI

Why it matters

Binary accuracy on reasoning benchmarks hides failures like prior collapse and inconsistency under paraphrase. This benchmark surfaces these pathologies and quantifies reasoning about parameter-dependent dynamics, guiding improvements in LLM reasoning capabilities.

Motivation

Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason about parameter-dependent dynamics.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.