Search papers, labs, and topics across Lattice.
This paper introduces SMART, a novel framework that leverages MLLM-generated motion descriptions to enhance temporal alignment for continuous sign language recognition (CSLR) and spotting under weak supervision. By integrating a Multi-Scale Temporal Adapter and a CSLR-guided spotting module (CSFormer), SMART effectively improves both recognition accuracy and temporal localization while operating efficiently with small batch sizes. Experimental results across four sign language benchmarks reveal significant performance gains, highlighting the framework's ability to unify recognition and spotting tasks in sign language processing.
SMART achieves significant improvements in sign language recognition and spotting by efficiently aligning video and text representations with minimal supervision.
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional video-text alignment also requires large batch sizes, making it inefficient for memory-intensive sign language video training. In this work, we propose SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting. SMART uses MLLMgenerated motion descriptions as auxiliary semantic cues and performs stable videotext alignment under small-batch training. To improve temporal representation learning, we introduce a Multi-Scale Temporal Adapter that models temporal interactions during transformer encoding. For dense temporal localization, SMART incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network. This unified framework enables CSLR features to benefit spotting, while spotting supervision complements weak CTC-based recognition. Experiments on four sign language benchmarks, including PHOENIX14-T, CSL-Daily, Large-scale KSL, and Disaster and Safety KSL datasets, demonstrate the effectiveness of SMART across both recognition and spotting tasks.