xPeer Benchmark: Simulating Scholarly Peer Review with Zero-Shot Reasoning
A recent preprint on arXiv (2510.02027) presents xPeer, a simulation engine designed for peer review, along with a benchmark consisting of two components to assess its effectiveness. The operational aspect examines 352 out of 500 simulation records that meet stable-task criteria across multiple fields and review types. The human-reference section provides 1,108 version-1 F1000Research manuscript records paired with human evaluations, employing a predetermined comparison of two humans against two xPeer outputs. Importantly, human-generated text and suggestions were not included in the input for generation, and source integration occurred only after the xPeer outputs were finalized, establishing a review withholding at the workflow level. Out of 802 records with two human evaluations, only 271 had two applicable xPeer reviewer fields, resulting in a complete-pair availability rate of 33.8%. This study seeks to enhance scalable evaluation in academic publishing, backed by verifiable evidence. The benchmark can be accessed through the xPeerd.com website. The paper outlines the methodology and initial results, highlighting the significance of zero-shot reasoning in simulating the peer review process, which is particularly pertinent given the rising volume of publications and the demand for efficient review mechanisms.
Key facts
- xPeer is a peer-review simulation engine delivered through xPeerd.com.
- The benchmark has two components: operational and human-reference.
- Operational component: 352 of 500 simulation records retained under stable-task criteria.
- Human-reference component: 1,108 F1000Research manuscript records with linked human reports.
- Comparison: two-human/two-xPeer design.
- Human-review text and recommendations were withheld from generation input.
- Complete-pair availability rate: 33.8% (271 out of 802 records).
- Paper available on arXiv: 2510.02027.
Entities
Institutions
- arXiv
- F1000Research
- xPeerd.com