Search papers, labs, and topics across Lattice.
This paper introduces MagicMakeup, a diffusion transformer framework that enhances regional controllability and fidelity in makeup transfer while preserving the source identity. By addressing pixel-to-attention misalignment and clarifying the separation between makeup transfer and identity preservation, the authors employ Token-Aligned Region Gating and Cross-Modal Perception Guidance to achieve precise and high-quality results. Extensive experiments demonstrate that MagicMakeup significantly outperforms existing methods in terms of controllability and fidelity across diverse styles and demographics.
MagicMakeup achieves unprecedented regional control in makeup transfer, ensuring high fidelity and identity preservation even across diverse styles and poses.
Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are (i) pixel-to-attention misalignment that causes spillover into non-target areas and weakens regional control; (ii) unclear transfer/preservation concept separation under two-image conditioning, leading to coupling between makeup attributes and identity; and (iii) the lack of a high-resolution dataset that is identity-consistent and region-labeled for fine-grained supervision. In this paper, we propose MagicMakeup, a diffusion transformer-based framework for region-controllable and high-fidelity makeup transfer, built on spatial constraints and concept disentanglement. To enable precise region-specific editing while preserving identity, we propose Token-Aligned Region Gating, which aligns pixel masks with attention and applies region-specific logit gating. To clarify the concepts of transfer and preservation, we further introduce Cross-Modal Perception Guidance, which aligns text and image features to enhance cross-modal concept perception. We also design a pipeline for the generation of 1024 x 1024 data pairs through region-specific makeup removal and establish a unified benchmark in synthetic and real settings. Extensive quantitative and qualitative experiments show that MagicMakeup improves regional controllability, makeup fidelity, and identity preservation, with strong robustness across styles, races, and poses.