Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-View Videos
Abstract
Instant reconstruction of dynamic humans and 3D human-object interactions from uncalibrated sparse-view videos is critical fornumerous downstream applications. Existing methods, however, are ei-ther limited by the slow reconstruction speeds or incapable of generat-ing novel-time representations. To address these challenges, we proposeForge4D, a feed-forward 4D human-object reconstruction and interpo-lation model that efficiently reconstructs temporally aligned represen-tations from uncalibrated sparse-view videos, enabling both novel-viewand novel-time synthesis. Our model simplifies the 4D reconstructionand interpolation problem as a joint task of streaming 3D Gaussian re-construction and dense motion prediction. For the task of streaming 3DGaussian reconstruction, we introduce learnable state tokens to enforcetemporal consistency in a memory-friendly manner. For novel-time syn-thesis, we design a novel motion prediction module to predict dense mo-tions for each 3D Gaussian between two adjacent frames. To overcomethe lack of the ground truth for dense motion supervision, we formulatedense motion prediction as a dense point matching task and introducea self-supervised retargeting loss to optimize this module. An additionalocclusion-aware optical flow loss is introduced to ensure motion consis-tency with plausible human and object movement, providing strongerregularization. Extensive experiments demonstrate the effectiveness ofour model on both in-domain and out-of-domain datasets. Code andmodels will be made publicly available.