Search papers, labs, and topics across Lattice.
This study critically examines the impact of speech preprocessing and dataset curation on Alzheimer's disease detection using the Pitt Corpus. The authors find that while speech-enhanced datasets improve performance on in-domain tasks, they significantly hinder model generalization across different datasets, leading to increased class imbalance and prediction shifts. Notably, matched enhancement during training and testing mitigates but does not fully resolve these issues, highlighting the complex trade-offs involved in dataset preparation for real-world applications.
Cleaner speech datasets may boost in-domain performance but can undermine generalization, revealing a hidden cost in Alzheimer's detection models.
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.