Search papers, labs, and topics across Lattice.
This paper introduces a novel neural compression method specifically designed for room impulse responses (RIRs) that incorporates structure-aware constraints to enhance acoustic reconstruction. By applying energy decay curve regularization and reverberant-speech supervision, the method effectively aligns the compression process with the unique characteristics of RIRs, addressing the limitations of existing audio codecs. Experimental results demonstrate that this approach achieves significantly lower reconstruction error and improved perceptual quality of reverberant speech at a low bitrate of 375 bps compared to traditional audio codecs.
Achieving superior RIR reconstruction and speech quality at just 375 bps challenges the effectiveness of conventional audio codecs in immersive audio applications.
Room impulse responses (RIRs) characterize the acoustic environment of a room by capturing how sound propagates and decays within an enclosed space. In applications such as immersive audio rendering, accurate acoustic reconstruction often relies on spatially densely sampled RIRs. This consequently gives rise to a large volume of RIR data, imposing a substantial burden on storage. Although recent neural audio codecs provide an effective framework for low-bitrate compression, their training objectives are mainly tailored to speech and general audio, and are therefore not well aligned with the acoustic characteristics of RIRs. Therefore, we propose an EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels. Specifically, at the RIR level, structure-aware constraints are imposed on the global decay behavior and local energy distribution of RIRs through energy decay curve (EDC) regularization and a short-time window energy constraint, while at the reverberant-speech level, reverberant-speech supervision is further introduced to constrain the consistency of the reverberant speech generated by the reconstructed RIRs. Experimental results show that, at a low bitrate of 375 bps, the proposed method achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.