Benchmark Radar
AI BENCHMARK PROFILE

FormalTCS

General AIMathematics & Formal Sciences

175 expert-validated instances from STOC, FOCS, SODA, and COLT papers (2025-2026) for evaluating LLMs on end-to-end theoretical computer science research, including autoformalization and proof tasks.

Released
2026-08-20
Readiness
Paper only
Primary field
General AI

Why it matters

Offers a realistic research evaluation for LLMs in theoretical computer science, highlighting bottlenecks like autoformalization and research taste.

Motivation

Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.