ARTFEED — Contemporary Art Intelligence

TCS-Bench: New Benchmark Evaluates LLMs on Theoretical Computer Science Proofs

ai-technology · 2026-08-15

A new benchmark called TCS-Bench has been developed by researchers to assess Large Language Models (LLMs) in generating proofs related to Theoretical Computer Science (TCS) at a research level. This benchmark features theorem-proving challenges sourced from leading theoretical computer science conferences, including STOC, FOCS, and SODA. Each challenge includes the context needed to create a complete proof for a specific result. The research tests cutting-edge models against this benchmark and confirms the accuracy of the generated proofs using a verification agent. The reference verifier demonstrates over 90% accuracy on a set labeled by experts, compared to human expert evaluations. This work can be found on arXiv with the identifier 2608.09538 in the Computer Science > Computation and Language section, aiming to enhance the evaluation of AI's formal reasoning and proof generation skills, which are vital for AI research and its applications in mathematics and computer science.

Key facts

  • TCS-Bench is a new benchmark for evaluating LLMs on theoretical computer science proof generation.
  • Tasks are derived from papers at STOC, FOCS, and SODA.
  • Each task provides context for a self-contained proof.
  • State-of-the-art models are evaluated on the benchmark.
  • A verification agent checks the correctness of generated proofs.
  • The reference verifier achieves over 90% accuracy on expert-labeled sets.
  • The verifier was benchmarked against human-expert proof judgments.
  • The paper is available on arXiv (2608.09538).

Entities

Institutions

  • arXiv
  • STOC
  • FOCS
  • SODA

Sources