Multimodal Pseudo-Labels Enhance Open-Vocabulary Segmentation
A recent study published on arXiv (2608.11681) introduces a multimodal approach aimed at improving open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS). This framework is designed to identify both known and novel object categories without the need for extensive annotations. Current techniques face difficulties with inaccurate pseudo-masks and insufficient visual-textual grounding. By leveraging pre-trained vision-language models, the framework facilitates automatic pseudo-label creation, CLIP-assisted synonym filtering, and GPT-driven caption reconstruction. It generates pseudo segmentation masks, relevant captions, and matched synonym sets using Grounded SAM, LLaVA, and CLIP, thereby enhancing visual-textual coherence. This advancement tackles a significant issue in computer vision, allowing models to recognize and segment objects beyond a static training dataset, which is vital for practical applications.
Key facts
- Paper arXiv:2608.11681 proposes a multimodal framework for open-vocabulary instance and panoptic segmentation.
- The framework uses pre-trained vision-language models for automatic pseudo-label generation.
- It employs CLIP-guided synonym filtering and GPT-based caption reconstruction.
- The method constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP.
- The target-vocabulary-assisted pseudo-labeling setting provides multimodal supervision without manual annotation.
- The work addresses challenges like noisy pseudo-masks, limited visual-textual grounding, and handling synonyms or out-of-vocabulary words.
- The paper is a cross-type announcement on arXiv.
- The approach enhances visual-textual alignment through three complementary components.
Entities
Institutions
- arXiv