The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
Zhuoyuan Wu ⋅ Xurui Yang ⋅ Jiahui Huang ⋅ Yue Wang ⋅ Jun Gao
Abstract
Estimating accurate camera poses, 3D scene geometry, and object mo-tion from in-the-wild videos is a long-standing challenge for classical structurefrom motion pipelines due to the presence of dynamic objects. Recent learning-based methods attempt to overcome this challenge by training motion estimatorsto filter dynamic objects and focus on the static background. However, their per-formance is largely limited by the availability of large-scale motion segmentationdatasets, resulting in inaccurate segmentation and, therefore, inferior structural 3Dunderstanding. In this work, we introduce the Dynamic Prior (D!"#$%) to robustlyidentify dynamic objects without task-specific training, leveraging the powerfulreasoning capabilities of Vision-Language Models (VLMs) and the fine-grainedspatial segmentation capacity of SAM2. D!"#$% can be seamlessly integrated intostate-of-the-art pipelines for camera pose optimization, depth reconstruction, and4D trajectory estimation. Extensive experiments on both synthetic and real-worldvideos demonstrate that D!"#$% not only achieves state-of-the-art performanceon motion segmentation, but also significantly improves accuracy and robustnessfor structural 3D understanding. Code is available.
Successful Page Load