ARTFEED — Contemporary Art Intelligence

Multimodal Pseudo-Labels Enhance Open-Vocabulary Segmentation

ai-technology · 2026-08-13

A recent study published on arXiv (2608.11681) introduces a multimodal approach aimed at improving open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS). This framework is designed to identify both known and novel object categories without the need for extensive annotations. Current techniques face difficulties with inaccurate pseudo-masks and insufficient visual-textual grounding. By leveraging pre-trained vision-language models, the framework facilitates automatic pseudo-label creation, CLIP-assisted synonym filtering, and GPT-driven caption reconstruction. It generates pseudo segmentation masks, relevant captions, and matched synonym sets using Grounded SAM, LLaVA, and CLIP, thereby enhancing visual-textual coherence. This advancement tackles a significant issue in computer vision, allowing models to recognize and segment objects beyond a static training dataset, which is vital for practical applications.

Key facts

  • Paper arXiv:2608.11681 proposes a multimodal framework for open-vocabulary instance and panoptic segmentation.
  • The framework uses pre-trained vision-language models for automatic pseudo-label generation.
  • It employs CLIP-guided synonym filtering and GPT-based caption reconstruction.
  • The method constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP.
  • The target-vocabulary-assisted pseudo-labeling setting provides multimodal supervision without manual annotation.
  • The work addresses challenges like noisy pseudo-masks, limited visual-textual grounding, and handling synonyms or out-of-vocabulary words.
  • The paper is a cross-type announcement on arXiv.
  • The approach enhances visual-textual alignment through three complementary components.

Entities

Institutions

  • arXiv

Sources