Articulated Object Reconstruction from Rest-State Observation
Abstract
Building interactive digital twins requires recovering both 3Dgeometry and the kinematic structures that govern how objects articu-late. Yet existing methods for articulated object reconstruction requireexplicitly observable motion from multiple articulation states. We intro-duce a rest-state formulation that reconstructs articulated objects froma single closed configuration, an inherently ill-posed setting where geom-etry, semantics, and motion priors compensate for the absence of motioncues. Our framework adopts an explicit mesh as an intermediate repre-sentation for cross-model verification and fusion, reconciling noisy out-puts from vision-language and segmentation models into spatially con-sistent part structures. To estimate joint parameters without observedmotion, we use a video diffusion model to synthesize articulation hy-potheses and validate them through geometric consistency. Our approachachieves accurate part decomposition and physically plausible articu-lation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.