Search papers, labs, and topics across Lattice.
This paper introduces a temporal-control methodology that enhances pretrained Diffusion Transformers (DiT) for video generation by enabling explicit time editing. The approach allows users to manipulate motion speed and temporal structure without the need to redesign the underlying model architecture. Key results indicate that the integration of a lightweight temporal module significantly expands the controllable dynamic range while maintaining the original generative capabilities of the DiT.
Unlocking precise control over motion dynamics in video generation could revolutionize how we create and edit video content.
Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics. We propose a temporal-control methodology that extends a pretrained DiT with explicit time editing, allowing control over motion speed and temporal structure without redesigning the backbone. Its core implementation augments the pretrained model with a lightweight temporal module, preserving the original generative prior while expanding its controllable dynamic range.