AI BENCHMARK PROFILE
TCS-BENCH
Evaluates LLMs on research-level theoretical computer science proof generation, using theorem-proving tasks from papers at STOC, FOCS, and SODA. Provides context for self-contained proofs and uses a verification agent to check correctness.
- Released
- 2026-08-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Fills a gap in evaluating LLMs on advanced formal reasoning in theoretical computer science, offering a structured protocol with a verification agent that aligns closely with expert judgment, supporting reliable model comparison.
Motivation
We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.