UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
Abstract
Fine-grained robotic manipulation depends on understand-ing parts, not only whole objects. Existing 3D foundation models tendto be either generalized but object-aware, or part-aware but limited toclosed-set taxonomies, which weakens zero-shot transfer. We study text-conditioned 3D part segmentation, where a free-form phrase selects afunctional part on point cloud. We introduce UniPart, a feed-forwardcross-modal 3D Transformer that conditions CLIP text embedding. Toscale supervision, we build LangPart-1M with 160K+ Objaverse assetsand 8M text to part pairs using multi-view consistent part generation.We further manually label a high-quality subset, LangPart-4K, for fine-tuning and evaluation. UniPart achieves strong zero-shot results on open-vocabulary part benchmarks and transfers to language-conditioned partgrasping in real world.