Search papers, labs, and topics across Lattice.
This paper introduces a privacy firewall pipeline that effectively removes speech from ambient audio while retaining essential environmental sounds for activity recognition in assisted living contexts. Utilizing a U-Net encoder-decoder trained on synthetic data, the model achieves 0% detectable speech across various conditions, significantly outperforming existing methods like Facebook Denoiser and ConvTasNet. The approach not only ensures speech privacy but also maintains high activity recognition accuracy, achieving 85% precision and recall post-speech removal, thus addressing a critical barrier to deploying acoustic sensing in elder care.
Achieving 0% detectable speech while preserving 85% activity recognition accuracy could revolutionize privacy in acoustic monitoring for elderly care.
Acoustic sensing offers a promising non-intrusive approach for monitoring daily activities of older adults, yet speech privacy concerns remain a critical barrier to real-world deployment. We present a privacy firewall pipeline based on a U-Net encoder-decoder, trained entirely on synthetic data, that removes speech from ambient audio while preserving environmental sounds indicative of daily activities. Activity recognition is performed using VGGish transfer learning with an SVM classifier. Evaluated on the ESC-50 and SINS datasets across multiple speech content levels, the proposed model reduced residual speech to 0% VAD-detectable speech (Silero Voice Activity Detection) under all tested conditions, outperforming Facebook Denoiser (6.55% residual), SepFormer (36.34%) and ConvTasNet (47.21%) on ESC-50 at the 100\% speech level. On ESC-50 at 40% speech level, classification performance recovers to 85% precision and 85% recall after speech removal, compared with 81%/75% before removal and an 84%/83% speech-free baseline. Evaluation on real-world participant home recordings collected with the AudioHive app showed 0% VAD-detectable speech after processing while maintaining 76% precision and recall. The pipeline enables privacy-preserving acoustic sensing without sacrificing activity recognition performance, addressing a key obstacle to the adoption of ambient monitoring in elderly care.