Search papers, labs, and topics across Lattice.
This paper establishes a theoretical foundation for Lipschitz-continuous neural networks specifically tailored for audio signal processing, addressing the limitations of existing architectures that process complex-valued signals. By introducing LipsAMs, a class of amplitude modifiers that operate solely on the magnitude of inputs, the authors derive necessary conditions for Lipschitz continuity and propose an efficient method for evaluating Lipschitz constants. The practical application of this framework is demonstrated through CoReM-LipsAM, a plug-and-play algorithm for audio signal recovery, which guarantees convergence and shows empirical success in speech dereverberation tasks.
Amplitude modifiers can ensure Lipschitz continuity in neural networks, enabling robust audio signal recovery with guaranteed convergence.
The Lipschitz continuity of deep neural networks (DNNs) is essential for establishing theoretical guarantees regarding their behavior. From both theoretical and practical perspectives, various methods have been proposed to construct Lipschitz-continuous architectures and control their Lipschitz constants. However, several DNN architectures common in audio signal processing fall outside the scope of existing theoretical frameworks, hindering the development of Lipschitz-continuous models in acoustic applications. In particular, despite their widespread adoption, DNNs that separately process the magnitude and phase of complex-valued signals cannot be Lipschitz continuous under existing frameworks. In this paper, to address this limitation, we establish a theoretical foundation for constructing amplitude modifiers (AMs), a class of DNN architectures that operate solely on the magnitude of a complex-valued input, with provable Lipschitz continuity. Specifically, we derive a necessary and sufficient condition for an AM to be Lipschitz continuous and propose LipsAMs (Lipschitz-continuous AMs) corresponding to common architectures for audio signals, including time-frequency masking. Furthermore, we develop an efficient framework for evaluating their Lipschitz constants and analytically derive these constants for some of the proposed architectures. As an application, we propose CoReM-LipsAM (Controlled Residual Maps via LipsAM) for plug-and-play (PnP) audio signal recovery, integrating a DNN as a data-driven prior within a model-based signal processing algorithm. The convergence of the obtained PnP algorithm is structurally guaranteed by the CoReM-LipsAM architecture and empirically validated through speech dereverberation experiments.