TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
2608.09538

Authors

Adarsh Kumarappan,Lalit Jain,Theophane Weber,Vincent Cohen-Addad,Dimitris Paparas

Abstract

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA).

Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark.

We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.

Resources

Ray graphicRay graphicRay graphicRay graphic

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.
Ray graphicRay graphicRay graphic