DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion
Abstract
Despite recent progress, direct text-driven 4D object gen-eration remains challenging yet highly desirable. In this paper, we in-troduce DiTex4D, a native text-to-4D generation framework that en-ables both text-driven 4D generation from scratch and 3D animationfrom static mesh. Built upon large-scale pre-trained 3D generation mod-els, our framework DiTex4D avoids intermediate text-to-video pipelinesand costly per-object optimization. Specifically, (i) we achieve 4D spa-tiotemporal consistency via inflating 3D attention with mixed-4D RoPEand tailored correlated noise injection strategy. (ii) To enable 3D anima-tion, we introduce a mask-based diffusion model conditioned on multi-view global context to maintain strict consistency with the initial frame.We further fine-tune the framework for 4D interpolation to synthesizehigh-frame-rate sequences with smoother motion. Extensive experimentsdemonstrate that DiTex4D can achieve higher-quality, semantically align-ed, and spatiotemporally coherent 4D object generation, surpassing mostexisting state-of-the-art text-to-4D generation methods.