Versatile Editing of Video Content, Actions, and Dynamics without Training
Abstract
Controlled video generation has seen drastic improvementsin recent years. However, editing actions and dynamic events, or insert-ing contents that should affect the behaviors of other objects in real-world videos, remains a major challenge. Existing trained models strug-gle with complex edits, likely due to the difficulty of collecting relevanttraining data. Similarly, existing training-free methods are inherently re-stricted to structure- and motion-preserving edits and do not supportmodification of motion or interactions. Here, we introduce DynaEdit,a training-free editing method that unlocks versatile video editing ca-pabilities with pretrained text-to-video flow models. Our method relieson the recently introduced inversion-free approach, which does not in-tervene in the model internals, and is thus model-agnostic. We showthat naively attempting to adapt this approach to general unconstrainedediting results in severe low-frequency misalignment and high-frequencyjitter. We explain the sources of these phenomena and introduce novelmechanisms for overcoming them. Through extensive experiments, weshow that DynaEdit achieves state-of-the-art results on complex text-based video editing tasks, including modifying actions, inserting objectsthat interact with the scene, and introducing global effects (see website).