Search papers, labs, and topics across Lattice.
This paper introduces the Autonomous Rectified Flow framework for generative speech enhancement, which eliminates the need for explicit time-step embeddings by leveraging a time-unconditional network. The authors demonstrate that the target vector field for denoising is inherently time-invariant, allowing for improved generation quality and robustness by focusing on spatial relationships rather than temporal conditioning. Key results show that this approach enhances inference efficiency and reduces overfitting to temporal trajectories, marking a significant advancement in generative speech enhancement techniques.
Time-unconditional generative speech enhancement achieves better quality and efficiency by sidestepping the constraints of temporal conditioning.
Most generative speech enhancement methods rely on explicit time-step embeddings for temporal conditioning. In this paper, we propose the Autonomous Rectified Flow framework, which challenges the necessity of such conditioning. Using a linear interpolation path, we show that the target vector field is inherently time-invariant. We further introduce a time-unconditional network that eliminates explicit time-step information and infers the denoising direction solely from the spatial relationship between the current state and the noisy observation. Predicting this target vector field is equivalent to modeling the noise distribution. By avoiding overfitting to temporal trajectories, the proposed autonomous design significantly improves generation quality, robustness, and inference efficiency.