Search papers, labs, and topics across Lattice.
This paper introduces ADEPS, a generative framework for Ambisonics encoding that integrates the physical acquisition model into the inference process to address hardware-dependent encoding artifacts caused by practical microphone arrays. By enabling zero-shot encoding across arbitrary array topologies, ADEPS offers a flexible solution that is not constrained by fixed array geometries, thus enhancing the spatial audio experience. Extensive evaluations reveal that ADEPS significantly surpasses traditional linear and parametric methods in both spatial fidelity and spectral quality, marking a substantial advancement in the field of spatial audio processing.
ADEPS can encode spatial audio with zero-shot flexibility across any microphone array, outperforming traditional methods in fidelity and quality.
Spatial audio enhances user immersion by reproducing 3D sound fields, with Ambisonics being a widely adopted representation. While Ambisonics is theoretically independent of the recording setup, practical microphone arrays introduce hardware-dependent encoding artifacts. Moreover, existing data-driven solutions lack flexibility, as they are typically restricted to fixed array geometries. To overcome these limitations, we propose ADEPS, a generative framework that explicitly embeds the physical acquisition model into the inference process. By leveraging this formulation, ADEPS effectively compensates for array-specific distortions while enabling zero-shot encoding across arbitrary array topologies. We train the underlying generative prior in an unsupervised manner solely on target Ambisonic representations. Extensive evaluations across diverse simulated and real microphone arrays demonstrate that ADEPS consistently outperforms both traditional linear and parametric baselines in spatial fidelity and spectral quality.