Search papers, labs, and topics across Lattice.
This paper conducts a variational-flow analysis of diffusion-based speech enhancement architectures, revealing that the sharp non-smooth transition in the SI-SDR degradation curve is localized to the predictor stage. The authors establish an exact factorization of parametric sensitivity, linking the behavior of the predictor output to the non-smoothness observed in the model's response to noise-power mismatch. By formulating hypotheses on the reverse-process flow, they provide a framework for understanding the conditions under which this non-smoothness occurs, setting the stage for future empirical validation.
The sharp kink in SI-SDR degradation reveals critical insights into the interplay between predictor outputs and noise-power mismatch in speech enhancement models.
Diffusion-based speech enhancement architectures that pair a deterministic predictor with a learned score network, exhibit a sharp non-smooth transition (``kink'') in the SI-SDR degradation curve at the training-time noise amplitude. We give a pathwise variational-flow analysis that localizes this non-smoothness to the predictor stage. The central identity is an exact factorization of the parametric sensitivity, $\partial \sig^{(M)} / \partial M = K(M) \cdot \partial C_M / \partial M$, where $K(M)$ is a continuous matrix-valued functional of the score Jacobian along the reverse trajectory and $C_M = 螤(y^{(M)})$ is the predictor output. Under three hypotheses on the reverse-process flow (score-Jacobian continuity, conditioning-Jacobian continuity, non-degeneracy of $K$), failure of $M \mapsto \sig^{(M)}$ to be $C^1$ at $M^\ast$ holds if and only if $M \mapsto 螤(y^{(M)})$ fails to be $C^1$ at $M^\ast$. We extend the localization to the finite-step Euler--Maruyama sampler actually run at inference. The hypotheses translate into a concrete experimental program; this paper specifies the program and presents the variational structure. The empirical validation is deferred to a companion experimental report.