Learning to Generate Rigid Body Interactions with Video Diffusion Models
Abstract
Recent video generation models have achieved remarkableprogress and are now deployed in film, social media production, and ad-vertising. Beyond their creative potential, such models also hold promiseas world simulators for robotics and embodied decision making. Despitestrong advances, current approaches still struggle to generate physicallyplausible object interactions and lack object-level control mechanisms. Toaddress these limitations, we introduce KineMask, an approach for videogeneration that enables realistic rigid body control, interactions, and ef-fects. Given a single image and a specified object velocity, our methodgenerates videos with inferred motions and future object interactions.We propose a two-stage training strategy that gradually removes futuremotion supervision via object masks. Using this strategy we train videodi!usion models (VDMs) on synthetic scenes of simple interactions anddemonstrate significant improvements and generalization to rigid bodyand hand-object interactions in real scenes. Furthermore, KineMask in-tegrates low-level motion control with high-level textual conditioning viapredicted scene descriptions, leading to support for synthesis of complexdynamical phenomena. Our experiments show that KineMask general-izes to di!erent VDMs and achieves strong improvements over recentmodels of comparable size. Ablation studies further highlight the com-plementary roles of low- and high-level conditioning in VDMs.Project Page: https://daromog.github.io/KineMask/