ARTFEED — Contemporary Art Intelligence

DinoSplat-OV: Training-Free Open-Vocabulary Segmentation for Remote Sensing

ai-technology · 2026-08-06

A new framework named DinoSplat-OV has been developed by researchers, designed to adapt the DINOv3 vision transformer for open-vocabulary semantic segmentation in remote sensing images without requiring training. This approach utilizes the recently released DINO.txt from DINOv3, which incorporates image-text contrastive learning into the DINO backbone, facilitating segmentation without the need for fine-tuning or pretraining. To tackle issues related to the dense distribution, multi-scale characteristics, and large dimensions of remote sensing imagery, DinoSplat-OV employs two main components: a Text-aware Laplacian Propagation module for enhancing patch-level predictions through textual and visual similarities, and a Gaussian Splatting Upsampling module for reconstructing pixel-level features. This framework is detailed in a paper available on arXiv (2608.03023) and aims to reduce the expenses associated with pixel-level annotations in remote sensing segmentation.

Key facts

  • DinoSplat-OV is a training-free framework for open-vocabulary semantic segmentation in remote sensing.
  • It adapts DINOv3, which includes DINO.txt for image-text contrastive learning.
  • The framework requires no fine-tuning or additional pretraining.
  • It targets dense distribution, multi-scale nature, and large size of remote sensing imagery.
  • Text-aware Laplacian Propagation module de-noises patch-level predictions.
  • Gaussian Splatting Upsampling module reconstructs pixel-level features.
  • The paper is available on arXiv with ID 2608.03023.
  • The method aims to reduce the need for costly pixel-level annotations.

Entities

Institutions

  • arXiv

Sources