EmoteGPT: 3D Human Facial Expression from Natural Language Descriptions
Abstract
Precise control of 3D facial expressions from text is crucialfor virtual avatars, animation, and human–computer interaction, yet ex-isting text-to-3D methods jointly generate identity, expression, and tex-ture, making fine-grained expression control difficult. We instead for-mulate text-driven expression synthesis as a regression problem in thedisentangled parameter space of a 3D Morphable Model (3DMM). Thissetting, however, requires paired data linking detailed language to pre-cise expression parameters, which is missing from existing resources. Tofill this gap, we introduce Txt2Emote, a benchmark of diverse 3D facialexpressions with fine-grained textual annotations obtained from GPT-4o and a high-fidelity face tracker, providing both explicit descriptionsdetailing facial features and implicit descriptions referencing the situa-tional context behind the expression. Leveraging this dataset, we presentEmoteGPT, a text-to-3D expression framework based on a Multi-ModalLarge Language Model (MLLM) with a dedicated