Search papers, labs, and topics across Lattice.
This paper introduces WeSep, a modular framework for Target Speaker Extraction (TSE) that reformulates the problem as a heterogeneous cue-conditioned learning task. By decoupling cue modules from separator backbones, WeSep allows for flexible integration of various auxiliary cues, facilitating systematic exploration of cue interactions and adaptability to diverse real-world scenarios. Experimental results show that WeSep maintains stable optimization across different modalities and cue availabilities, highlighting its potential for enhancing TSE performance in practical applications.
WeSep reveals that decoupling cue modules from separator architectures can significantly enhance the adaptability and performance of Target Speaker Extraction systems in real-world environments.
The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSep, a unified framework that reformulates TSE as a heterogeneous cue-conditioned learning problem. In WeSep, cue modules and separator backbones are decoupled through standardized interfaces, enabling configurable cue injection and flexible integration of diverse modalities. The design enables systematic study of cue structure, intra- and cross-modal interaction, and dynamic cue availability within a shared optimization framework, facilitating adaptation to real-world conditions. Experiments across enrollment, spatial, visual, and textual cues reveal modality-dependent characteristics and demonstrate stable optimization under heterogeneous cue availability. The toolkit will be publicly available.