Search papers, labs, and topics across Lattice.
This paper introduces ToPO, a novel method for optimizing token-conditioned preferences in attention-based latent diffusion models, which constructs a spatial-temporal routing mechanism to enhance image generation. By leveraging preferred-branch cross-attention and an auxiliary pixel-midpoint ordering term, ToPO outperforms the existing Diffusion-DPO approach across multiple metrics, including SD-1.5 and HPSv2. The results indicate that ToPO achieves higher endpoint estimates and better performance in blind A/B studies, suggesting significant improvements in the efficiency and quality of image generation processes.
ToPO achieves superior image generation performance by optimizing token-conditioned preferences, outperforming existing methods across multiple evaluation metrics.
Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from branchwise squared-residual contrast in a frozen reference denoiser. Preferred-branch cross-attention uses content tokens to modulate the spatial factor, and an auxiliary pixel-midpoint ordering term is added without local labels or a learned reward model. In matched three-seed retrainings with a shared update schedule, ToPO has higher endpoint estimates than Diffusion-DPO on all five reported SD-1.5 metrics and on HPSv2, ImageReward, and CLIP for SDXL. It also receives larger raw win shares in an aggregate blind SDXL A/B study. These findings are scoped to the reported equal-update U-Net protocols rather than an equal-compute comparison.