Search papers, labs, and topics across Lattice.
This paper introduces appearance pointers, a novel mechanism for enhancing Diffusion Transformers (DiTs) by enabling precise regional control over image generation through the alignment of text and image inputs with user-defined masks. By employing a region correspondence network and a spatial aggregation mechanism, the authors achieve effective multimodal guidance without the need for extensive retraining, allowing for multiple regional descriptions while maintaining a manageable token load. The results demonstrate that their approach not only matches but often exceeds the performance of existing modality-specific methods, marking a significant advancement in controllable image synthesis.
Appearance pointers enable unprecedented localized control in image generation, outperforming traditional methods without the need for retraining.
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.