OEIS Open: Benchmark Tests AI's Ability to Prove Mathematical Conjectures
Researchers have introduced OEIS Open, a benchmark designed to evaluate the ability of language models (LMs) to solve open mathematical conjectures. The benchmark is based on 492 open conjectures from the Online Encyclopedia of Integer Sequences (OEIS), formalized in the Lean proof assistant by Tsoukalas and colleagues. Unlike previous attempts that used a specialized agent, OEIS Open provides open-source evaluation code that can run any generic LM, with safeguards against cheating. The study found that LMs equipped with a minimal set of tools resolved 147 of the conjectures (30%) with a budget of $50 per attempt. A smaller subset, OEIS Open Lite, consisting of 100 randomly selected conjectures, allows for cheaper evaluation. With a budget of $200 per attempt, the best current LM achieved a 44% success rate on OEIS Open Lite. Interestingly, providing LMs access to 476,000 papers from arXiv did not improve performance, nor did using more sophisticated agent loops. The conjectures in the benchmark are of uncertain mathematical significance, and most have likely received little prior attention. The work is detailed in a paper on arXiv (2608.11941) and aims to provide a standardized way to measure progress in automated theorem proving.
Key facts
- OEIS Open is a benchmark based on 492 open mathematical conjectures from the OEIS.
- The conjectures are formalized in Lean by Tsoukalas et al.
- The evaluation code is open-source and runs any generic language model.
- LMs with minimal tools resolved 147 conjectures (30%) with a $50 budget per attempt.
- OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation.
- With a $200 budget, the best LM scored 44% on OEIS Open Lite.
- Access to 476,000 arXiv papers did not improve performance.
- More sophisticated agent loops also did not improve performance.
- The conjectures are of uncertain mathematical significance and likely have received little attention.
Entities
Institutions
- OEIS
- Lean
- arXiv