Search papers, labs, and topics across Lattice.
This paper introduces MTVDiff, a multimodal latent diffusion framework designed to tackle the challenges of thermal-to-visible face translation, such as geometric discontinuities and identity degradation. By integrating depth and textual information through innovative modules like the Dual-Branch Cross-Attention Fusion and Gated Text-to-Visual Feature Alignment, MTVDiff significantly enhances image quality and face verification performance. Experimental results on the MCXFace and SpeakingFaces datasets show that MTVDiff achieves up to 48.3% reduction in FID and 8.9% improvement in Rank-1 accuracy compared to existing methods.
Achieving up to 48.3% better image quality and 8.9% higher accuracy in face recognition, MTVDiff revolutionizes thermal-to-visible face translation through advanced multimodal integration.
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.