Search papers, labs, and topics across Lattice.
This paper identifies and formalizes the issue of manifold drift in preference optimization for generative models, which occurs when reward-driven updates cause terminal samples to deviate from the pretrained data manifold. The authors propose ThermoDPO, a temperature-controlled objective that effectively anchors pairwise preference optimization to preferred samples, thereby mitigating the risk of manifold drift. Experimental results demonstrate that ThermoDPO-weighted significantly outperforms existing methods, achieving a StrictScore of 0.899 and improving performance metrics on the SD3.5-M benchmark by up to 47.5%.
Manifold drift can lead to substantial misalignment in generative models, but ThermoDPO offers a powerful solution that anchors preference optimization to the pretrained data manifold.
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers the terminal data distribution, whereas a preference update leaves the pretrained manifold whenever its induced terminal displacement has a nonzero normal component. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO and controls a pointwise reconstruction-based surrogate for manifold distance. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted. On the main toy benchmark, ThermoDPO-weighted attains a StrictScore of 0.899, compared with 0.629 for FlowDPO and 0.857 for FlowDPO+RFT. On SD3.5-M at CFG = 4.5, it improves OCR by 47.5% and the average of four metrics by 16.0%.