MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
Abstract
Generating lifelike facial animation for dyadic conversationsrequires reconciling high-level cognitive intent with precise low-level mo-tor reflexes, yet existing methods fall short in the semantic understandingof dialogue context and in precise dynamic control. In this paper, we pro-pose MindFlow, a dual-pathway generative framework inspired by theVentral-Dorsal pathway model in neuroscience, which decouples genera-tion into two collaborative streams, thereby harmonizing deep semanticreasoning with fine-grained control. In the Ventral module, we trans-form the conventional Sentence-Action approach into a novel Chunk-State approach that models raw acoustic streams as a context-aware,evolving emotional state chain, capturing subtle paralinguistic nuancesand mid-utterance emotional shifts missed by sentence-level modeling.The Dorsal module features a conditional autoregressive flow matchingnetwork for high-fidelity facial motion, driven by high-frequency acousticcues and modulated by emotion states, plus a Selective Acoustic Injectorfor adaptive audio gating to ensure robustness in talking-and-listeningdynamics without interference. Extensive experiments demonstrate thatMindFlow achieves superior semantic appropriateness and motion natu-ralness compared to state-of-the-art baselines.