Search papers, labs, and topics across Lattice.
This study benchmarks various upsampling strategies for Latent Acoustic Mapping (LAM), a self-supervised method for generating high-resolution acoustic maps from multichannel recordings. The research highlights the significant degradation of LAM's performance with sparse 4-channel arrays and evaluates how different upsampling architectures can mitigate this issue. Key findings reveal that while the original full-resolution LAM outperforms others, lightweight models trained separately yield competitive results, emphasizing the importance of representation alignment over model complexity.
Sparse microphone arrays can dramatically hinder acoustic mapping performance, but the right upsampling strategy can bridge the gap.
Latent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.