Search papers, labs, and topics across Lattice.
This paper introduces DREAM, a novel multimodal framework designed to enhance video retrieval through dual-objective encoding that improves both visual and textual representation. By employing a hybrid language modeling strategy and a hierarchical vision encoder with cascaded group attention, DREAM effectively captures fine-grained temporal dependencies and complex linguistic structures. The model achieves state-of-the-art performance on benchmark datasets MSRVTT, MSVD, and LSMDC, demonstrating significant improvements in retrieval accuracy and contextual coherence.
Achieving new state-of-the-art retrieval scores, DREAM shows that dual-objective encoding can significantly enhance cross-modal alignment in video content.
In today's media-driven world, the exponential growth of video content across domains such as surveillance, education, and entertainment has made retrieving semantically relevant videos via natural language queries increasingly critical. Early video retrieval systems relied on handcrafted features or shallow cross-modal mappings, limiting their ability to capture complex semantics and temporal dynamics. While large-scale vision-language models have improved cross-modal alignment, challenges remain in modeling fine-grained temporal dependencies and nuanced linguistic structures. In this paper, we introduce DREAM: Dual-path Representation Enhancement and Alignment Model, a novel multimodal framework that addresses these limitations through enhanced visual and textual encoding. DREAM incorporates a hybrid language modeling strategy that combines masked and permuted language modeling objectives to capture both local and global linguistic semantics. On the visual side, we design a hierarchical vision encoder with cascaded group attention, which integrates spatial and temporal information through multi-stage token interaction and coarse-to-fine attention refinement. We validate DREAM through comprehensive evaluations on the widely-used MSRVTT, MSVD and LSMDC benchmark datasets, where it achieves new state-of-the-art R1 scores of 49.4%, 49.7% and 27.3%, respectively. Qualitative analyses further show the model's ability to maintain coherent attention across frames and align complex queries with dynamic video content. These findings underscore the effectiveness of hierarchical attention and dual-objective textual modeling in enabling robust, context-aware video retrieval, and pave the way for future research in advancing cross-modal representation learning.