Search papers, labs, and topics across Lattice.
This paper addresses the issue of Prior Collapse in Test-Time Tuning (TTT) for pretrained diffusion models used in video editing, where models often discard essential text conditions and spatial latents. The authors introduce ElasticTTT, a framework that incorporates Target Distribution Regularization, Contrastive CFG, and an Asynchronous Noise Schedule to maintain the generative prior and enhance model flexibility. Their extensive evaluations show that ElasticTTT achieves state-of-the-art performance in one-shot video editing, effectively mitigating the foundational mismatch between generative models and standard TTT approaches.
Prior Collapse in video editing models can be overcome, leading to state-of-the-art one-shot editing performance without sacrificing generative quality.
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.