Search papers, labs, and topics across Lattice.
This study conducts a layer-wise probing analysis of a transformer ASR encoder to investigate how dysarthric speech affects internal representations across different conditions, including original dysarthric speech and TTS resynthesis. The results reveal a hierarchy of representation where phoneme boundary information remains weak throughout the layers, while phoneme identity becomes more recoverable in the upper layers, indicating that disordered speech significantly impacts high-level representations. Additionally, the research demonstrates that targeted adaptations in specific layers can achieve performance close to full encoder adaptation, highlighting the potential for efficient fine-tuning in low-resource dysarthric ASR tasks.
Disordered speech alters high-level representations in ASR models, revealing that phoneme identity is recoverable only in the upper layers, which has significant implications for model adaptation strategies.
Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We present a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. The probes reveal a task-dependent hierarchy: phoneme boundary information stays weak for dysarthric speech at every layer, phoneme identity becomes recoverable toward the upper layers, and recognition difficulty is encoded in the deepest layers. Tone-sensitive evaluation shows Mandarin lexical tone is a persistent error source. Cross-condition similarity divergence grows with depth, indicating that disordered speech affects high-level representations more than low-level acoustic features. Guided by these findings, single-layer LoRA at layer 7 and adaptation on subset layers 5-8 achieve performance within 3.5% and 2.48% relative margins of full encoder adaptation, respectively, while upper-layer adaptation is less effective for dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.