LLMs struggle most with converting informal mathematical claims into formal statements—not with proving them—suggesting that bridging the gap between natural language mathematics and formal verification is the key challenge for autonomous mathematical research.
This paper introduces FormalTCS, a benchmark of 175 frontier theoretical computer science research problems from top venues, with expert-verified formal proofs in Lean.