Benchmark Radar
AI BENCHMARK PROFILE

LCS-Bench

General AIMathematics & Formal Sciences

LCS-Bench is a theory-scale benchmark for auto-formalization in logics for computer science. It includes 327 textbook items, over 4,076 Lean declarations, and supports five evaluation tracks with definitional equivalence checkers.

Released
2026-06-25
Readiness
Paper only
Primary field
General AI

Why it matters

Auto-formalization at theory scale remains challenging. LCS-Bench provides a reusable benchmark to measure consistency, faithfulness, and correctness of formalization systems, enabling progress in scalable verification.

Motivation

Auto-formalization is critical for scalable formal verification, but existing progress largely focuses on isolated statements, while theory-scale auto-formalization, which coherently translates hundreds of interdependent definitions, lemmas, and theorems, remains open due to challenges in consistency, faithfulness, scalability, and correctness.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.