ChaosBench-Logic
ChaosBench-Logic v2 evaluates LLM logical reasoning over dynamical systems with 40,886 questions across 165 systems, 27 first-order logic predicates, and 78 axiom edges. The CARE protocol measures calibration and adversarial robustness, reporting metrics such as MCC.
- Released
- 2026-05-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Binary accuracy on reasoning benchmarks hides failures like prior collapse and inconsistency under paraphrase. This benchmark surfaces these pathologies and quantifies reasoning about parameter-dependent dynamics, guiding improvements in LLM reasoning capabilities.
Motivation
Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason about parameter-dependent dynamics.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.