Search papers, labs, and topics across Lattice.
This paper introduces VisGate, a framework for multimodal sequential recommendation that adaptively fuses visual and collaborative signals based on item embeddings and user interaction history. By treating visual utility as a contextual variable rather than a fixed property, the authors demonstrate that visual information's contribution to recommendation quality varies significantly across items and user contexts. The findings reveal that visual utility is particularly beneficial in scenarios of interaction sparsity, providing insights into when and why visual features enhance recommendation performance.
Visual utility in recommendations isn't static; it fluctuates based on user context and item characteristics, revealing critical insights into effective multimodal fusion.
Multimodal sequential recommender systems commonly fuse visual and collaborative signals uniformly, treating visual features as generically informative regardless of item or user context. We argue that visual utility, defined as the contribution of visual signals to recommendation quality, is a latent contextual variable that depends on both the item and the user's interaction history rather than a fixed item property. To model this variability, we introduce VisGate, a framework that makes adaptive item-level fusion decisions conditioned on item embeddings and the user's current sequence context. Visual representations are learned through a contrastive objective over sequential co-occurrence patterns, preserving complementarity with collaborative embeddings rather than aligning them into a shared space. Beyond achieving competitive recommendation performance, VisGate's learned gate serves as a measurement tool for understanding when and why visual information is beneficial. Our analyses show that visual utility varies across items, increases under interaction sparsity when collaborative signals are weak, and correlates with visual distinctiveness in semantically meaningful ways. Together, these findings highlight the importance of both fine-grained fusion and modality complementarity, while demonstrating that item-level visual utility can be estimated and interpreted through learned gating behaviour.