PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
Abstract
In this work, we focus on photorealistic sign avatar modeling,which is crucial for effective communication with the Deaf community andis characterized by complex hand gestures and nuanced facial expressions.To this end, we introduce MVSign, the first multi-view Chinese signlanguage dataset co-designed with Deaf experts, featuring diverse gesturesand rich annotations. For precise SMPL-X annotation, we develop a hybridfitting pipeline that produces accurate body, hand, and facial parametersand can also be applied to the monocular setting. Building on MVSign,we propose a decoupled sign avatar representation that isolates body,head, and hand components to capture complex articulations, togetherwith a motion-aware sampling strategy to handle motion blur and balancegesture diversity. Extensive experiments demonstrate that our methodachieves high-fidelity visual results on MVSign, particularly in detailedhand and facial regions, and generalizes well to in-the-wild monocularsign language videos. Project page: https://naaapi.github.io/PHOSA.