InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
Abstract
Recent approaches in controllable novel view video generationoften rely on fine-tuning pre-trained Video Di!usion Models (VDMs).This dominant paradigm is computationally expensive and frequentlysu!ers from catastrophic forgetting of the model’s original generative pri-ors. To address this challenge, here we propose InverseCrafter, a VDMtraining-free framework that reformulates novel view video generationas an inpainting-based inverse problem in the latent space, eliminatingthe need for any annotated 4D training data. The core of our methodis to establish operator equivalence by employing a lightweight latentmask encoder to define a latent-domain masking operation via a con-tinuous, multi-channel representation. This principled representationfaithfully models the forward process in the latent domain, enablinge"cient, backpropagation-free solvers while bypassing the costly bottle-neck of repeated VAE operations. InverseCrafter achieves high-fidelity,spatio-temporally coherent novel view synthesis with near-zero additionalinference overhead and excels at general-purpose video inpainting andediting by fully preserving the pre-trained VDM’s generative capabilities.