SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
Abstract
Diffusion Transformers (DiTs) have significantly advancedaudio-driven portrait animation, but their high computational cost leadsto substantial inference latency. Although training-free diffusion cachingaccelerates inference significant, existing methods are primarily devel-oped for text-conditioned generation and overlook the spatial and modal-ity imbalances inherent in audio-driven portrait animation. In this pa-per, we propose SyncCache, a training-free caching acceleration methodtailored for DiT-based portrait animation that explicitly exploits asym-metric dynamics. Specifically, high-frequency dynamics driven by audioconditions and concentrated in human regions are more challenging andcritical to cache and reuse than the low-frequency visual background inportrait animation. First, we introduce Spatially-Asymmetric Probingto prioritize error sensitivity in dynamic human region. Second, throughModality-Decoupled Caching, we bypass heavy DiT block by reusingstable inter-block residuals, while continuously recomputing lightweightaudio blocks to preserve precise lip synchronization. Furthermore, weintroduce a cache ratio to control cache capacity and formulate memory-adaptive cache selection as an offline dynamic programming problemwithout online overhead. Extensive experiments demonstrate that Sync-Cache achieves superior speed–quality trade-offs, delivering up to 4.12×acceleration on HunyuanVideo-Avatar and 3.75× on Wan-S2V with near-lossless visual fidelity and precise audio alignment.