Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to phone segmentation and recognition by leveraging self-supervised speech models (S3Ms) through a method called S3M-based Phonological Activation Mapping (SPAM). By mapping S3M representation frames to phonological feature activations and employing lightweight prediction heads, the authors demonstrate that their method can effectively perform both tasks with minimal phonetic transcription data. The results show strong performance across various datasets, indicating that phonetic structure can be efficiently extracted from existing S3M representations without extensive retraining.
Phonological features can be extracted from self-supervised speech models in under a minute, achieving state-of-the-art performance in phone segmentation and recognition.
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.