LLM-Based Peer Review: Official Guidelines Outperform Imitation
Researchers in computer science have conducted a study examining the impact of various reviewer guidelines on automated peer review utilizing LLMs. The results indicate that official conference guidelines align most closely with human evaluations, whereas guidelines that mimic reviewers based on high-quality human feedback tend to be less effective. Additionally, the implementation of rigid rubric-style scoring consistently hampers performance, underscoring the importance of subjective and comprehensive assessment methods. These findings imply that evaluation criteria developed through conference practices can effectively guide automated review processes.
Key facts
- Peer review automation is increasingly necessary due to growing workload.
- Official conference guidelines yield review results most consistent with human judgments.
- Reviewer-imitating guidelines generated from human reviews using LLMs are less effective.
- Strict rubric-style scoring consistently degrades performance.
- Subjective and holistic scoring is important for effective automated review.
- The study is published on arXiv under Computer Science > Computation and Language.
- The paper is titled 'Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review'.
- The arXiv ID is 2607.22553.
Entities
Institutions
- arXiv