Search papers, labs, and topics across Lattice.
This paper introduces HybridSB-MoE, a dual-domain framework for generative speech enhancement that addresses the limitations of existing spectral and waveform models by leveraging asymmetric uncertainty fusion and heterogeneous expert routing. The method integrates epistemic uncertainty from spectral paths and aleatoric variance from waveform bridges, allowing for adaptive mixing weights that cater to different error regimes. Experimental results on the VoiceBank+DEMAND dataset demonstrate that HybridSB-MoE outperforms both diffusion and Schr枚dinger Bridge-based baselines while maintaining competitive performance with few-step consistency-distilled methods.
HybridSB-MoE achieves superior speech enhancement by intelligently fusing spectral and waveform models, outperforming traditional methods in both efficiency and quality.
Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schr\"odinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.