Skip to yearly menu bar Skip to main content


Poster

DNI: Dilutional Noise Initialization for Diffusion Video Editing

Sunjae Yoon · Gwanhyeong Koo · Ji Woo Hong · Chang D. Yoo

Strong blind review: This paper was not made available on public preprint services during the review process Strong Double Blind
[ ]
Fri 4 Oct 1:30 a.m. PDT — 3:30 a.m. PDT

Abstract:

Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation stems from an initial latent noise employed in diffusion video editing systems. The diffusion video editing systems prepare initial latent noise to edit by gradually infusing Gaussian noise onto the input video. However, we observed that the visual structure of the input video still persists within this initial latent noise, thereby restricting non-rigid editing such as motion change necessitating structural modifications. To this end, this paper proposes Dilutional Noise Initialization (DNI) framework which enables editing systems to perform precise and dynamic modification including non-rigid editing. DNI introduces a concept of 'noise dilution' which adds further noise to the latent noise in the region to be edited to soften the structural rigidity imposed by the input video, resulting in more effective edits closer to the target prompt. Extensive experiments demonstrate the adaptability and effectiveness of the DNI framework. The code will be made publicly available.

Live content is unavailable. Log in and register to view live content