DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
Abstract
Video generation models achieve high visual quality but of-ten struggle to generate physics-aware videos. Unlike rigid-body motion,which can be described by explicit trajectories or formulas, complex de-formation dynamics remain challenging to synthesize. We observe that alack of physical reasoning for localizing dynamic areas allows irrelevantregions to dilute the model’s attention, leading to generation failure.In this paper, we propose DeforM, a reasoning-guided image-to-videogeneration framework that directs the model’s focus toward physics-critical regions. To reason and localize these critical regions, we introducea VLM-guided physical reasoning module, DeforM-Reason, to identifytarget objects and generate spatial-temporal masks. For physical guid-ance, we develop two alternative strategies: DeforM-Free for training-freemechanism analysis, and DeforM-Injection as a powerful training-basedgenerator. Experimental results demonstrate that DeforM improves therealism of generated deformation scenarios, outperforming baseline mod-els in both visual quality and physical consistency.