CORE-3D: Context-Aware Open-Vocabulary Retrieval in 3D Scenes
The research article 'CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D' (arXiv ID: 2509.24528) focuses on retrieving objects from 3D environments. Current zero-shot, open-vocabulary techniques rely on embedding vectors derived from vision-language models (VLMs) to create 2D class-agnostic masks, which often results in fragmented masks and incorrect semantic assignments. To tackle this issue, the authors employ SemanticSAM with a progressive granularity refinement approach to generate better object-level masks, minimizing over-segmentation that occurs in models like vanilla SAM, thus enhancing 3D semantic segmentation. Additionally, they introduce a context-aware CLIP encoding method that integrates various contextual views for each mask. The paper is under revision and pertains to computer vision, 3D scene comprehension, and embodied AI, with implications for robotics and augmented reality.
Key facts
- Paper title: CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- Published on arXiv with ID 2509.24528
- Announce type: replace-cross
- Addresses object retrieval from 3D scenes
- Existing methods use VLMs to generate 2D masks and project them into 3D
- Problems: fragmented masks and inaccurate semantic assignments
- Proposes using SemanticSAM with progressive granularity refinement
- Improves downstream 3D semantic segmentation
- Employs context-aware CLIP encoding with multiple contextual views
- Aims to improve open-vocabulary 3D semantic mapping
Entities
Institutions
- arXiv