ARTFEED — Contemporary Art Intelligence

LLM Reviews Align with Human Peer Review at ICLR 2026

ai-technology · 2026-08-06

A recent study uploaded on arXiv, titled 2608.03659, evaluates the performance of three prominent large language models (LLMs) during peer reviews of 300 submissions to the 2026 ICLR conference, categorized into oral, poster, and rejected. The models under scrutiny are OpenAI's GPT-5.4, Google’s Gemini 3.1 Pro Preview, and Anthropic’s Claude Opus 4.6. Researchers assessed their alignment with decision categories, variations in recommendations, and their ability to identify weaknesses. While the models effectively distinguished between accepted and rejected works, they encountered challenges differentiating oral from poster presentations, shedding light on their capabilities and limitations in academic peer reviews.

Key facts

  • Study compares OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6
  • 300 topic-matched ICLR 2026 submissions used
  • Papers equally divided among oral, poster, and rejected
  • Models used identical instructions and rating scales
  • Decision information removed before review
  • All three LLMs distinguished accepted from rejected papers
  • None reproduced the oral versus poster distinction
  • Study analyzes alignment with decisions, recommendation-scale usage, and thematic agreement

Entities

Institutions

  • OpenAI
  • Google
  • Anthropic
  • ICLR

Sources