Search papers, labs, and topics across Lattice.
This paper introduces IT-TextFusion, an iterative framework for text-guided image fusion that enhances multi-modal integration by employing deep Cross-Attention and multi-scale Cross-Gate Fusion. By leveraging text-conditioned feature interaction across multiple stages, the method effectively addresses complex image degradations while maintaining the integrity of both visible and infrared information. Experimental results demonstrate significant improvements in information preservation and perceptual quality metrics compared to existing methods, highlighting the framework's robustness and adaptability to various datasets.
Iterative text-guided image fusion can significantly enhance perceptual quality and information preservation, even in the presence of complex degradations.
Text-guided image fusion has recently emerged as an effective paradigm for integrating multi-modal information while enabling flexible and task-oriented fusion control. However, existing text-guided fusion methods often rely on shallow semantic-visual interaction and limited attention mechanisms, which restrict their ability to robustly handle complex degradations and fully exploit textual guidance. In this paper, we propose an iterative text-guided image fusion framework that incorporates text-conditioned feature interaction across multiple fusion and refinement stages. The proposed method integrates deepest-level Cross-Attention, multi-scale Cross-Gate Fusion, and stage-specific text-conditioned modulation, allowing the global text embedding to condition hierarchical feature fusion and residual refinement. By repeatedly injecting the pooled text embedding across hierarchical decoder and refinement stages, the proposed framework provides degradation-aware global semantic conditioning while preserving complementary information from the visible and infrared modalities. Experiments on several benchmark datasets show that the proposed method improves several information-preservation and perceptual-quality metrics, while exhibiting metric-dependent trade-offs on some datasets.