DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
Abstract
While 2D dix001Busion models have achieved remarkable successin identity-preserving personalization, extending this capability to 3Dassets remains a signix001Ccant challenge due to the complexities of multi-view consistency and spatial control. Inspired by these 2D advance-ments, we present a novel personalization method for text-guided 3Dediting that enables compositional, object-level control through naturallanguage. Given a 3D input, we render orthogonal views and extractobject-level segmentation masks to isolate semantic components. Wethen learn distinct token embeddings for each component through a tai-lored two-phase optimization strategy: multi-view textual inversion withattention alignment, followed by full x001Cne-tuning of multi-view dix001Busionmodel. During inference, these disentangled tokens seamlessly composewith editing prompts to generate multi-view consistent images, whichare subsequently lifted into high-x001Cdelity textured 3D meshes. Extensiveevaluations across diverse editing scenarios demonstrate that our methodsuccessfully transfers the x001Dexibility of 2D personalization to 3D, achiev-ing state-of-the-art edit faithfulness and identity preservation comparedto existing baselines.