Search papers, labs, and topics across Lattice.
This paper introduces two novel spectrally weighted STFT loss functions aimed at improving lightweight streaming speech enhancement by mitigating magnitude over-attenuation in mid-to-high frequency regions. The proposed sigmoid-weighted loss and signal-dependent spectrally adaptive loss enhance phase-aware contributions and are evaluated using the HyST-Net architecture, which employs hybrid MHA-GRU spectral-temporal modeling for low-latency applications. Experimental results demonstrate that these losses lead to significant improvements in high-frequency spectral reconstruction, with the spectrally adaptive loss further optimizing mid-frequency performance for a balanced output across the frequency spectrum.
Spectrally adaptive loss functions can drastically improve high-frequency speech clarity in streaming applications, addressing a critical gap in current enhancement techniques.
This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency regions caused by the magnitude-phase compensation effect. The proposed sigmoid-weighted loss applies a smooth frequency-dependent modulation to the phase-aware contribution, while the signal-dependent spectrally adaptive loss further conditions the modulation on the ground-truth log-magnitude spectrogram. To evaluate the proposed objectives, we additionally design HyST-Net, a lightweight and competitive backbone with hybrid MHA-GRU spectral-temporal modelling for low-latency streaming scenarios. Experimental results exhibit consistent improvements in high-frequency spectral reconstruction for both losses. The spectrally adaptive loss further enhances the mid-frequency region, resulting in a more balanced spectral reconstruction across the full frequency range.