Search papers, labs, and topics across Lattice.
This study addresses the challenges of automatic phoneme recognition in non-canonical speech by introducing a linguistically structured approach that decomposes phoneme predictions into articulatory feature dimensions such as manner, place, and voicing. Utilizing a hierarchical multi-task learning architecture, the method employs task-specific heads and a cross-attention fusion module to enhance phoneme prediction accuracy. The results demonstrate significant performance improvements over baseline models, highlighting the effectiveness of articulatory feature supervision for robust and interpretable phoneme recognition in pathological speech contexts.
Articulatory feature supervision can dramatically enhance phoneme recognition in non-canonical speech, revealing interpretable error patterns that align with phonological structures.
Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme recognition systems are typically trained on canonical speech and treat phonemes as atomic categorical labels, limiting their ability to detect structured articulatory errors common in speech disorders and accents. In this work, we introduce a linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. We implement this formulation using a hierarchical multi-task learning architecture in which task-specific articulatory feature heads learn feature-level representations that are subsequently integrated through a cross-attention-based fusion module to produce phoneme predictions. To address the scarcity and noise of pathological speech labels, we combine this framework with semi-supervised learning via Momentum Pseudo-Labeling (MPL) and propose a cascaded training strategy that progressively introduces articulatory feature tasks while employing staged unfreezing of a pretrained speech encoder. Experiments on L2-ARCTIC, used as a proxy for pathological speech variation, show that the proposed approach achieves substantial improvements in phoneme recognition performance compared to strong baseline architectures, while yielding interpretable error patterns aligned with phonological feature structure. These results suggest that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech, and motivate future validation on clinically diagnosed pathological speech datasets.