EndoVLM: Anatomy-Guided Vision-Language Model for Endoscopy
A new vision-language foundation model called EndoVLM has been developed for analyzing endoscopic images. This innovative model has been pre-trained on more than 348,000 endoscopic procedures, linking each clinical report with its related image set. EndoVLM utilizes an Anatomy-Guided Sparse Pooling approach, where textual descriptions serve as queries to enhance sparse attention, allowing for the effective consolidation of semantically relevant frames into anatomy-specific visual representations from overlapping image sets. Following this, a Progressive Semantic-Aware alignment process is implemented. This research tackles the disparity between structured anatomical descriptions and uncurated visual data, aiming to harness the rich semantic insights found in clinical reports. The study can be accessed on arXiv with the identifier 2608.04472.
Key facts
- EndoVLM is a vision-language foundation model for endoscopy.
- Pre-trained on over 348,000 endoscopic examinations.
- Each examination pairs a clinical report with an image collection.
- Uses Anatomy-Guided Sparse Pooling to aggregate salient frames.
- Progressive Semantic-Aware alignment is employed.
- Addresses modality gap between textual descriptions and visual streams.
- Paper available on arXiv (2608.04472).
Entities
Institutions
- arXiv