Search papers, labs, and topics across Lattice.
This study introduces specialized phonetic forced alignment models for Chengdu Mandarin, addressing the gap in alignment systems for low-resource language varieties. By training a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuning a pretrained audio encoder for text-independent alignment (Chengdu-FC), the authors achieved significant improvements over Standard Mandarin baselines. The Chengdu-MFA model reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a remarkable 61.2% reduction, paving the way for more accurate aligners in under-resourced languages.
Chengdu Mandarin aligners outperform Standard Mandarin systems, achieving up to 61.2% improvement in phonetic accuracy.
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.