Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding
Abstract
The ability to remotely measure an animal’s size and shape(i.e. its morphology) in its natural habitat is of key importance for ap-plications such as conservation, re-identification, and biomechanical un-derstanding. Monocular depth estimation (MDE) and 3D reconstructionapproaches can be used. This is useful to generate accurate deformablemesh models or to characterize biologically important parameters such assexual dimorphism or the size distribution of populations. Whilst depthestimation and 3D reconstruction have been extensively studied as coretopics in computer vision, ranging from early work on simple, rigid ob-jects, to more recent work on human and animal reconstruction, themajority of existing animal models are trained on video/image baseddatasets that lack metric scale and ground truth e.g. from camera trapimages or from public videos. In addition, they do not fully reflect thechallenges of accurately estimating the focal animal’s scale when it is dis-tant from a camera. To address this limitation, we present WildDepth,a multimodal dataset and benchmark suite for depth estimation, behav-ior detection, and 3D reconstruction from diverse categories of animalsranging from domestic to wild environments with synchronized RGB andLiDAR. We provide three focused benchmarks: (1) monocular depth es-timation with per-distance and temporal stability analysis, (2) behav-ior detection, (3) 3D reconstruction and densification. Our results showthat large-scale MDE models degrade significantly at long range, whileLiDAR-anchored fusion reduces metric error by up to 25% – 30% RMSEin mid-range scenarios. Our aim is to enable camera trap data to bemore accurately ‘lifted’ to 3D morphometrics, unlocking a new era ofscale-aware animal modelling.⋆ M. Aamir, N. Muramatsu, and S. Shin contributed equally and are listed alpha-betically by surname.