Search papers, labs, and topics across Lattice.
The paper introduces MSA-EchoLite, a novel lightweight acoustic echo cancellation (AEC) framework that employs an asymmetric dual-branch encoder and an echo-aware frequency-time modulation (EAM) module to enhance performance without significantly increasing computational costs. By addressing the limitations of existing systems that rely on downsampling, MSA-EchoLite effectively models the discrepancies and correlations between microphone inputs and echo-related features. Experimental results indicate that this Bark-domain variant achieves 99.1% of the perceptual evaluation of speech quality (PESQ) score of its frequency-domain counterpart while requiring only 26.1% additional FLOPs, thus demonstrating a superior performance-complexity trade-off.
MSA-EchoLite achieves near state-of-the-art performance in acoustic echo cancellation with a fraction of the computational cost, redefining efficiency in lightweight AEC systems.
Existing lightweight acoustic echo cancellation (AEC) systems often combine linear AEC with Bark-domain DNN-based suppression to lower the computational footprint. In such systems, downsampling layers further compress the input features into a compact bottleneck representation, but this compression weakens frequency-time modeling capacity and degrades performance. To mitigate this limitation, we propose MSA-EchoLite, a lightweight Bark-domain AEC framework with an asymmetric dual-branch encoder and an echo-aware frequency-time modulation (EAM) module. The EAM module enriches the compressed bottleneck representation by modeling discrepancy and correlation cues between the dual-branch microphone and echo-related latent features. Experimental results show that the Bark-domain variant of MSA-EchoLite offers a better performance-complexity trade-off than its frequency-domain counterpart but is more sensitive to feature compression. With only 26.1% additional FLOPs over its non-EAM Bark-domain variant, its EAM-enhanced version achieves 99.1% of the PESQ of the frequency-domain counterpart, which requires nearly twice the FLOPs, and even surpasses it in SDR. Overall, MSA-EchoLite outperforms state-of-the-art lightweight AEC models while using only 0.2 M parameters and 100 M FLOPs/s.