Search papers, labs, and topics across Lattice.
This study introduces language orthogonalization, a method that mitigates the language identity bias in self-supervised speech models (S3Ms) to enhance cross-lingual detection of Parkinson's disease (PD). By applying a closed-form ridge residualization against language embeddings derived from healthy-control speech, the authors effectively decouple language-specific features from pathology-related variations. The results demonstrate significant improvements in PD detection across multiple languages and tasks, addressing the critical issue of high specificity coupled with low sensitivity in existing classifiers.
Language orthogonalization transforms self-supervised speech models, enabling them to detect Parkinson's disease across languages without being misled by language identity.
Self-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure.