Search papers, labs, and topics across Lattice.
This paper investigates zero-shot Parkinson's Disease (PD) detection from speech using both handcrafted acoustic features analyzed by a general-purpose LLM and raw audio waveforms processed by audio-capable models. They evaluate these approaches across four languages, finding that handcrafted features offer more stable performance in low-resource languages, while raw audio input can provide dataset-dependent improvements. The results demonstrate that the choice of input modality significantly impacts zero-shot PD detection performance.
Turns out, how you feed speech data to AI for Parkinson's detection鈥攈andcrafted features versus raw audio鈥攄rastically changes its accuracy, especially in different languages.
Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted from speech or the raw audio waveform itself, affects performance for Parkinson's disease (PD) detection across different languages. In this study, we systematically compare two input modalities for zero-shot PD detection: (i) handcrafted acoustic features extracted from speech recordings analyzed by a general-purpose LLM, and (ii) direct waveform input analyzed by audio-capable models. Experiments on PD speech datasets in four languages show that performance varies across input modalities, speech tasks, and languages. Handcrafted acoustic features provide more stable performance in a low-resource language (e.g., Bengali), whereas audio input yields dataset-dependent gains. These findings highlight the impact of input modality on zero-shot PD detection from speech.