AI BENCHMARK PROFILE
FormalTCS
175 expert-validated instances from STOC, FOCS, SODA, and COLT papers (2025-2026) for evaluating LLMs on end-to-end theoretical computer science research, including autoformalization and proof tasks.
- Released
- 2026-08-20
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Offers a realistic research evaluation for LLMs in theoretical computer science, highlighting bottlenecks like autoformalization and research taste.
Motivation
Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.