Search papers, labs, and topics across Lattice.
This paper introduces a multi-level framework for dementia detection from speech that addresses privacy concerns by neutralizing eavesdropping threats at both the signal and feature levels. Utilizing a Cumulative Signal Attack (CSA) to induce transcription errors while preserving prosodic features, and a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection to suppress speaker identity, the method achieves near-chance speaker identification. The results demonstrate strong dementia classification performance, with an F1 score of 0.78 and an AUC of 0.86, highlighting the effectiveness of the approach in balancing privacy and utility.
Near-chance speaker identification is achieved while maintaining strong dementia classification, showcasing a breakthrough in privacy-preserving AI applications.
Speech recordings used for dementia detection inherently expose speaker identity, raising critical privacy concerns. Existing methods typically address only singular threats and fail to resolve the privacy--utility trade-off. We propose a multi-level framework designed to neutralize two distinct eavesdropping vectors. At the signal level, a Cumulative Signal Attack (CSA) concentrates perturbations in keyword-aligned regions to maximize transcription error (Word Error Rate WER = 1.00) while preserving vital prosodic biomarkers. At the feature level, a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection suppresses speaker-discriminative dimensions while retaining dementia-relevant diagnostic structure. Evaluated on the DementiaBank Pitt Corpus, our framework achieves near-chance speaker identification (Equal Error Rate EER = 0.59, F1 = 0.003) while maintaining strong dementia classification performance (F1 = 0.78, AUC = 0.86).