AI Chatbots' Study Retrieval Quality for Clinical Questions Evaluated
A study recently shared on arXiv (2608.13786) looked into how effectively three large language model chatbots—Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5—retrieve clinical research. The researchers crafted clinical questions based on 20 review queries from the 2026 Cochrane Database of Systematic Reviews, role-playing as patients, clinicians, and evidence-synthesis researchers. Each chatbot was tested four times for each role, resulting in a total of 720 responses. They were asked to support their answers with primary clinical citations, which were then measured against the Cochrane reviews' datasets. This study highlights the importance of assessing the quality of studies retrieved by these chatbots, especially as their use in clinical settings grows.
Key facts
- Study evaluates three LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5.
- Clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews.
- Simulated user roles: patient, clinician, and evidence-synthesis researcher.
- Each chatbot queried under each user role with four independent repetitions, yielding 720 responses.
- Chatbots asked to support answers with primary clinical citations.
- Citations benchmarked against included and excluded study sets of the Cochrane reviews.
- On average, a chatbot response retrieved a certain number of relevant studies (exact number not provided in abstract).
- Study addresses gap in evaluating quality of retrieved studies and factors driving selection.
- Prior research focused on citation fabrication.
- Study published on arXiv with identifier 2608.13786.
Entities
Institutions
- arXiv
- Cochrane Database of Systematic Reviews