AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
Abstract
We present AHOY, a method for reconstructing complete,animatable 3D Gaussian avatars from in-the-wild monocular video de-spite heavy occlusion. Existing methods assume unoccluded input: afully visible subject, often in a canonical pose. This excludes the vastmajority of real-world footage where people are routinely occluded by fur-niture, objects, or other people. Reconstructing from such footage posesfundamental challenges: large body regions may never be observed, andmulti-view supervision per pose is unavailable. We address these challengeswith four contributions: (i) a hallucination-as-supervision pipeline thatuses identity-finetuned diffusion models to generate dense supervisionfor previously unobserved body regions; (ii) a two-stage canonical-to-pose-dependent architecture that bootstraps from sparse observations tofull pose-dependent Gaussian maps; (iii) a map-pose/LBS-pose decou-pling that absorbs multi-view inconsistencies from the generated data;(iv) a head/body split supervision strategy that preserves facial iden-tity. We evaluate on YouTube videos and on multi-view capture datawith significant occlusion and demonstrate state-of-the-art reconstruc-tion quality. We also demonstrate that the resulting avatars are robustenough to be animated with novel poses and composited into 3DGSscenes captured using cell-phone video. Our project page is available athttps://miraymen.github.io/ahoy/.