Study Evaluates Multimodal LLMs for Optical Diagnosis of Colorectal Polyps
A retrospective diagnostic performance study evaluated the accuracy of five multimodal large language models (MLLMs) in classifying colorectal polyps and predicting histology from endoscopic images. The models—Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5—were tested on the PRIME dataset, a curated collection of white light and narrow-band imaging (NBI) images. For each model, F1 scores, percent correct scores, and accuracy were calculated for Paris classification, Narrow-band Imaging Colorectal Endoscopic (NICE) classification, and predicted histology, compared against expert responses for 132 cases. Statistical significance was assessed using Cochran's Q and McNemar's Test. Results indicated that all models achieved F1 scores above 0.9 for distinguishing neoplastic versus non-neoplastic polyps, suggesting high diagnostic accuracy. The study, announced on arXiv with identifier 2608.07543, highlights the potential of MLLMs to assist in optical diagnosis, which could guide resection strategies and surveillance decisions. The findings underscore the growing role of artificial intelligence in medical imaging and diagnostic support.
Key facts
- Study evaluated five MLLMs: Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, GPT-5.
- Used PRIME dataset with white light and NBI images.
- 132 cases were analyzed.
- F1 scores >0.9 for all models for neoplastic vs. non-neoplastic.
- Statistical tests: Cochran's Q and McNemar's Test.
- Study is retrospective diagnostic performance study.
- Published on arXiv (2608.07543).
- Potential to guide resection strategy and surveillance.
Entities
Institutions
- arXiv