Search papers, labs, and topics across Lattice.
This paper introduces a unified approach to acoustic scene analysis for hearing aids by utilizing a causal Deep Neural Network to estimate time- and frequency-dependent power proportions of speech, music, and noise from acoustic mixtures. By integrating multiple tasks into a single framework, the method reduces computational complexity and enhances the interpretability of the scene representation. The proposed system demonstrates performance on Voice Activity Detection that rivals state-of-the-art estimators while offering a more comprehensive description of the listening environment.
A unified DNN-based approach reveals richer acoustic scene descriptions while maintaining competitive performance in Voice Activity Detection.
Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.