Search papers, labs, and topics across Lattice.
This study introduces the Spatial Audio Representation Learning (SARL) benchmark to systematically evaluate the spatial encoding capabilities of pretrained audio models. By probing both source-level and room-level factors, the authors uncover that spatial encoding is significantly influenced by input configuration and training paradigms, with source factors being more easily decoded than room factors. The findings highlight systematic biases in existing audio representations, providing a foundation for improved evaluation and development of spatial audio models.
Spatial audio models reveal a surprising bias: source factors are decoded with far greater accuracy than room characteristics, challenging assumptions about their representational capabilities.
Pretrained spatial audio encoders are increasingly used as general-purpose representations for perceptual tasks, yet their spatial encoding capabilities remain poorly understood. We introduce the Spatial Audio Representation Learning (SARL) benchmark, a controlled framework for evaluating spatial information in pretrained audio models. SARL probes source-level factors (azimuth, elevation, distance, class) and room-level factors (RT60, volume, shape). Experiments across diverse encoders reveal three patterns: input configuration and training paradigm shape spatial encoding; source factors are consistently easier to decode than room factors; and sensitivity analysis under controlled perturbations shows heterogeneous responses to source and room variation. These results reveal systematic biases in current pretrained audio representations. SARL is released as an open-source benchmark for reproducible evaluation of spatial audio representations.