Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Abstract
Diffusion Transformers (DiTs) have advanced video gener-ation with high-quality, temporally coherent results. However, extend-ing them to motion transfer, which requires following reference motionwhile aligning with a target prompt, remains challenging due to lim-ited understanding of motion and structure representations within DiTs.We analyze video DiTs at the attention-head level and identify distinctheads specialized for motion and spatial structure. Based on this insight,we propose a head-aware controllable motion transfer framework thatrequires no parameter updates. Our method refines motion cues frommotion-specialized heads via semantic correspondence guidance and pre-serves structure through selective feature injection. This head-level con-trol not only enables accurate motion transfer but also provides an in-terpretable foundation for controllable video generation with DiTs.