REVEAL: Largest Generative Foundation Model for Endoscopy Trained on 5 Million Frames
A groundbreaking generative foundation model for endoscopy, named REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), has been unveiled by researchers. This model utilizes GastroNet-5M (GN-5M), a comprehensive dataset featuring 5 million endoscopic images from various centers. REVEAL effectively bridges the divide between natural and clinical imagery while mitigating the high computational demands of training large Diffusion Transformers by utilizing encoders specifically pretrained on endoscopic data. This technique aligns diffusion latents with visual features unique to the domain, ensuring the retention of detailed textures and complex anatomical elements. Additionally, REVEAL functions as a robust feature extractor, achieving impressive results across several benchmarks. Further details can be found in a paper available on arXiv (arXiv:2608.07176v1).
Key facts
- REVEAL is the largest generative foundation model for endoscopy.
- Trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames.
- Uses encoders pretrained on endoscopic distribution to align diffusion latents.
- Preserves fine textures and anatomical structures.
- Serves as a feature extractor with strong benchmark performance.
- Paper available on arXiv with ID 2608.07176v1.
Entities
Institutions
- arXiv