Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
Abstract
Conversational talking face generation has recently attractedincreasing attention, aiming to synthesize interactive talking videos wherecharacters speak, listen, and respond dynamically to each other. Thistask presents three core challenges: 1) Flexibility: enabling multi-rounddialogues with an arbitrary number of participants; 2) Naturalness: main-taining coherent motion and appropriate non-verbal feedback through-out the interaction; and 3) Efficiency: achieving real-time generation andlow computation overhead for long-term continuous online conversation.Despite recent advances, existing methods still fall short in balancingall three requirements. To bridge this gap, we introduce InterTalk, anovel and efficient framework designed for highly interactive conver-sational talking face generation. Built upon a motion-based architec-ture, InterTalk supports real-time conversation synthesis. Our methodachieves strong flexibility by explicitly modeling multi-round conversa-tional dynamics among each participant, eliminating constraints on theirnumbers. To enhance interactivity, we incorporate motion feedback frommultiple participants and introduce an iterative generation strategy formore natural behaviors. Besides, we disentangle motion into several fa-cial components, enabling targeted refinements for natural response suchas precise lip-sync and realistic eye-blinking. Finally, we construct a newmulti-person conversational dataset and enrich it with 3D face-baseddata augmentation. Extensive experiments demonstrate that InterTalkachieves superior interaction quality while maintaining real-time perfor-mance at 30 FPS.