DinoSplat-OV: Training-Free Open-Vocabulary Segmentation for Remote Sensing
A new framework named DinoSplat-OV has been developed by researchers, designed to adapt the DINOv3 vision transformer for open-vocabulary semantic segmentation in remote sensing images without requiring training. This approach utilizes the recently released DINO.txt from DINOv3, which incorporates image-text contrastive learning into the DINO backbone, facilitating segmentation without the need for fine-tuning or pretraining. To tackle issues related to the dense distribution, multi-scale characteristics, and large dimensions of remote sensing imagery, DinoSplat-OV employs two main components: a Text-aware Laplacian Propagation module for enhancing patch-level predictions through textual and visual similarities, and a Gaussian Splatting Upsampling module for reconstructing pixel-level features. This framework is detailed in a paper available on arXiv (2608.03023) and aims to reduce the expenses associated with pixel-level annotations in remote sensing segmentation.
Key facts
- DinoSplat-OV is a training-free framework for open-vocabulary semantic segmentation in remote sensing.
- It adapts DINOv3, which includes DINO.txt for image-text contrastive learning.
- The framework requires no fine-tuning or additional pretraining.
- It targets dense distribution, multi-scale nature, and large size of remote sensing imagery.
- Text-aware Laplacian Propagation module de-noises patch-level predictions.
- Gaussian Splatting Upsampling module reconstructs pixel-level features.
- The paper is available on arXiv with ID 2608.03023.
- The method aims to reduce the need for costly pixel-level annotations.
Entities
Institutions
- arXiv