Reconstruction Benchmark Tests AI's Ability to Recover Research Ideas from Bibliographies
A new evaluation standard known as Reconstruction challenges language models to extract the fundamental research concept of a published paper solely from its pre-publication references, excluding the paper itself and any related literature. This benchmark implements a rigorous anti-leakage strategy, featuring temporal citation cutoffs, anonymous reference identifiers, and static bibliographies to avoid prompt-time leakage of the original idea. In six scientific fields and across 643 assessed papers, seven leading models achieved only modest Match rates of around 3-15%. Subsequently, a reference-only multi-agent system, which integrated cross-model assessments with a Swiss tournament format for aligned hypothesis slots, was tested without external web searches, raising Match rates to about 20%. This indicates that collaborative reasoning among multiple models can enhance idea recovery. Introduced in a paper on arXiv (ID: 2608.16645), this benchmark offers a fresh perspective on evaluating AI's capability to deduce underlying research ideas from bibliographic data alone. The results underscore the current limitations of language models in this area, while also indicating that multi-agent collaboration could improve outcomes. The benchmark's framework effectively addresses data leakage concerns and creates a controlled setting for hypothesis generation and evaluation.
Key facts
- Benchmark named Reconstruction
- Withholds seed paper and contemporaneous/future literature
- Uses temporal citation cutoff, anonymous reference IDs, frozen bibliographies
- Evaluated on 643 papers across six scientific domains
- Seven frontier models achieved Match rates of 3-15%
- Multi-agent pipeline with cross-model review and Swiss tournament
- Pipeline raised Match rates to approximately 20%
- No external web search used in pipeline
- Paper available on arXiv (ID: 2608.16645)
Entities
Institutions
- arXiv