Search papers, labs, and topics across Lattice.
This paper introduces a Geometry-Aware Dynamic Convolution (Geo-DConv) framework that enhances multi-channel speech enhancement (SE) systems by incorporating explicit microphone geometry information, allowing for robust performance across varying array configurations. The approach addresses the limitations of existing array-agnostic SE methods, which often overlook critical spatial filtering cues provided by microphone placements. Experimental results on the RealMAN dataset show significant performance improvements for fixed-array models when adapted to array-invariant settings, highlighting the effectiveness of leveraging geometric priors in SE tasks.
By integrating microphone geometry into speech enhancement, this framework achieves consistent performance gains across diverse array configurations, challenging the status quo of fixed-array limitations.
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.