- Updated: August 25, 2026
- 1 min read
CORE-3D: Context‑aware Open‑vocabulary Retrieval by Embeddings in 3D – A Breakthrough in 3D Scene Understanding
CORE-3D: Context‑aware Open‑vocabulary Retrieval by Embeddings in 3D
We are excited to present CORE-3D, a novel framework that pushes the boundaries of zero‑shot, open‑vocabulary 3D semantic mapping. By integrating SemanticSAM with progressive granularity refinement and a context‑aware CLIP encoding strategy, CORE-3D delivers significantly more accurate object‑level masks and richer visual context for downstream 3D tasks.
Key Innovations
- Enhanced Mask Generation: Leveraging SemanticSAM reduces over‑segmentation and produces high‑quality, object‑centric masks compared to vanilla SAM.
- Context‑aware CLIP Encoding: Multiple contextual views of each mask are weighted empirically, providing a deeper semantic understanding.
- Zero‑Shot Retrieval: Enables language‑driven object retrieval without task‑specific training.
Performance Highlights
Extensive experiments on benchmark datasets demonstrate notable improvements in 3D semantic segmentation and language‑based object retrieval, outperforming existing state‑of‑the‑art methods.
Read the Full Paper
Explore the complete study on arXiv and discover how CORE-3D can transform your 3D perception pipelines.
Explore More on ubos.tech
Stay tuned for upcoming releases and implementation details.
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.