EMOTE: Expressive Motion and Shape Disentanglement for Human Animation
Abstract
High-fidelity and expressive controllable human animation isessential for content creation and digital avatar applications. However,existing methods face a dilemma between expressiveness and disentangle-ment. Mainstream 2D pose-conditioned approaches suffer from "motion-shape entanglement", leading to the leakage of the driving subject’s bodyshape. Conversely, methods relying on 3D priors (e.g., SMPL) achievegeometric disentanglement but struggle to capture facial expressions andcomplex gestures, resulting in rigid animations. To this end, we proposeEMOSH, a novel framework for high-fidelity controllable human videogeneration. First, an Expressive Human Model (EHM) is introduced asthe core control representation. By explicitly disentangling shape andpose parameters, we fundamentally resolve the body shape leakage issue.Alongside this, a robust motion tracker is designed to accurately estimateEHM parameters from video. Second, we propose a Coarse-to-Fine Hy-brid Motion Injection strategy, enabling more fine-grained control overexpressions and gestures. Furthermore, we introduce a Spatially-AlignedConditioning mechanism to bridge the domain gap between training and∗ †Intern at WeChat Vision. Corresponding authors.inference, improving identity consistency. Extensive experiments demon-strate that EMOSH outperforms previous methods in both self-drivenand cross-driven scenarios, producing high-fidelity videos with vivid ex-pressions while maintaining shape disentanglement. Video demos andadditional results are available at our Project page.