Search papers, labs, and topics across Lattice.
This paper introduces an array-generic neural framework for direction-of-arrival (DoA) estimation that leverages complex directional array transfer functions (ATFs) to enhance generalization across different microphone arrays. By employing separate convolutional encoders for multichannel spectrograms and ATF metadata, the method utilizes cross-attention to effectively fuse representations and predict source directions in a Cartesian vector format. Experimental results demonstrate that the approach maintains performance across various unseen array configurations, including those resembling mobile devices, outperforming traditional and learning-based methods in challenging acoustic environments.
Generalizing DoA estimation across diverse microphone arrays could revolutionize audio processing in mobile and dynamic environments.
Direction-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during training and generalize poorly to unseen devices. This paper proposes an array-generic neural DoA estimation framework using measured or simulated complex directional array transfer functions (ATFs) matched to real-world multi-microphone devices. The method processes multichannel spectrograms and ATF metadata with separate convolutional encoders, fuses the resulting representations through cross-attention, and predicts source directions using a multi-source Cartesian vector output formulation. Experiments on simulated 2D and 3D localization tasks under reverberation and diffuse babble noise show that the proposed approach generalizes to previously unseen arrays, including mobile-phone-like configurations, without major performance degradation, while remaining competitive with conventional and learning-based baselines.