Search papers, labs, and topics across Lattice.
This paper introduces DiffM2A, a geometry-adaptive conditional diffusion framework designed for robust Ambisonic encoding from sparse microphone arrays with variable topologies. By employing a Geometry-Adaptive Spherical Harmonic Projection (GASHP) to create boundary-aware steering functions and utilizing a dual-branch Elucidated Diffusion Model, the method effectively addresses the challenges of noise amplification and overfitting in higher-order Ambisonic encoding. Evaluations reveal that DiffM2A significantly outperforms traditional and neural baseline methods in terms of signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation, even with unseen microphone layouts and varying boundary conditions.
DiffM2A achieves superior Ambisonic encoding fidelity, maintaining performance across diverse microphone configurations and boundary conditions.
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.