Search papers, labs, and topics across Lattice.
This paper investigates the attention mechanisms within Diffusion Transformers (DiTs) to enhance motion transfer capabilities in video generation. By analyzing attention heads, the authors identify specialized heads for motion and spatial structure, leading to a novel framework that allows for head-aware control of motion transfer without requiring parameter updates. The proposed method effectively refines motion cues and preserves spatial structure, resulting in improved accuracy and interpretability in controllable video generation.
Unlocking head-level control in Diffusion Transformers enables precise motion transfer without the need for retraining, revolutionizing video generation capabilities.
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.