Search papers, labs, and topics across Lattice.
This paper introduces DreOPD, a novel method that enhances flow-matching models by integrating degraded-reference extrapolative on-policy distillation to optimize task-specific rewards more effectively. By converting implicit reward extrapolation into closed-form velocity regression, DreOPD achieves the stability of on-policy distillation while avoiding the high-variance gradients typical of trajectory-level optimization. Experimental results demonstrate that DreOPD not only surpasses traditional on-policy distillation and multi-task reinforcement learning baselines but also outperforms specialized teachers across various metrics.
DreOPD achieves superior performance by transforming reward extrapolation into stable velocity regression, outperforming traditional methods and specialized teachers alike.
Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.