Search papers, labs, and topics across Lattice.
This study introduces a novel method for estimating speaker head orientation using the phase component of the short-time Fourier transform from a single microphone array, processed through a deep neural network architecture that integrates convolutional, recurrent, and self-attention layers. The approach outperforms traditional methods reliant on handcrafted features or raw audio inputs, achieving state-of-the-art accuracy on a large-scale dataset that includes both simulated and real-world recordings. Notably, the model demonstrates significant improvements in personalization, achieving a mean angular error of just 11.3 degrees when fine-tuned to individual users and environments.
Achieving a mean angular error of just 11.3 degrees in head orientation estimation could revolutionize applications in smart environments and driver monitoring.
Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverages the phase component of the short-time Fourier transform from a single microphone array as input to a deep neural network combining convolutional, recurrent, and self-attention layers. Unlike prior methods that use physics-informed handcrafted features or raw waveform inputs, our approach enables robust learning from simulated and real data. Trained on a large-scale dataset generated with voice directivity patterns and fine-tuned on real recordings, our model achieves state-of-the-art accuracy, outperforming baselines under both clean and noisy conditions. Personalization experiments further demonstrate significant gains, reaching a mean angular error of 11.3 degrees when adapting to individual users and environments.