Learning Video Dynamics with Predictive Differentiable Rendering
Abstract
How to accurately predict a high-fidelity future world? Whilethe visual world is inherently continuous, existing deterministic videoprediction models operate in discrete pixel space and are mainly opti-mized with pixel-wise mean squared error (MSE), which often leads toover-smoothed predictions and a lack of fine-grained visual details. Toaddress these limitations, we propose Predictive Differentiable Rendering(PDR), a novel end-to-end video prediction paradigm that bridges thegap between discrete and continuous representations. Inspired by recentprogress in 3D reconstruction with 3D Gaussian Splatting, we introducePredGS, a lightweight and plug-and-play adapter based on 2D Gaus-sian representation, which could be seamlessly integrated with existingpixel space predictors, significantly improving spatial detail preserva-tion with negligible computational overhead. Furthermore, we developpredgsplat, a CUDA-accelerated differentiable 2D Gaussian renderersupporting arbitrary channels. Each Gaussian is defined by 5 + C learn-able parameters (position, scale, rotation, and C channel amplitudes)and achieves up to 10× faster rendering than the baseline. Optimizedby a combined L1 and SSIM loss, PDR overcomes the inherent blurringtendencies of MSE Loss, significantly enhancing the prediction perfor-mance. Extensive experiments on diverse real-world benchmarks, includ-ing TaxiBJ, WeatherBench, KTH, and Human3.6M, demonstrate thatPDR consistently surpasses existing methods, delivering superior detailpreservation, visual fidelity, and predictive accuracy.