Search papers, labs, and topics across Lattice.
This study investigates the efficacy of self-supervised learning (SSL) speech representations for detecting Parkinson's disease (PD) across different languages and datasets. By conducting a layer-wise analysis of nine SSL backbones with logistic regression probes, the authors demonstrate that the optimal representation layer is largely determined by the source dataset rather than the architecture itself. Additionally, the findings reveal that classifiers trained on PD data exhibit a lack of pathological specificity, assigning similar probabilities to PD and dementia speech, which underscores significant limitations for clinical application.
Layer selection for speech-based PD detection is more about the dataset than the model architecture, revealing a critical flaw in current approaches.
Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related characteristics or exploit dataset-specific confounds, particularly since most SSL backbones are pretrained exclusively on healthy speech. To investigate this question, we perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages. We structure the evaluation as multiple scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology. Our results reveal two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus. These results highlight critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.