Search papers, labs, and topics across Lattice.
This paper details the SonicAGI submission to the REAL-TSE Challenge, addressing the complexities of real-world target speaker extraction (TSE) by employing a data-centric approach that integrates simulated and real meeting audio. The authors introduce SwiftNet-Lookahead, which optimizes latency while maintaining performance, and USEF-TFGridNet, which enhances speaker fidelity through advanced enrollment cross-attention mechanisms. The results, with SwiftNet-Lookahead placing second in Track 1 and USEF-TFGridNet fifth in Track 2, demonstrate the effectiveness of real-data-oriented training and tailored modeling strategies for improving conversational TSE outcomes.
Real-data-oriented training combined with innovative modeling strategies led to competitive placements in the REAL-TSE Challenge, showcasing significant advancements in target speaker extraction.
Real-world target speaker extraction (TSE) remains challenging because target speech, interference, and enrollment are recorded under mismatched acoustic conditions with reverberation, noise, and irregular conversational overlap. This paper describes the SonicAGI submission to the REAL-TSE Challenge (IEEE SLT 2026). We take a data-centric approach that combines fully simulated mixtures from clean speech with real meeting overlaps, and use a frozen offline enhancer to provide a denoised mirror of real targets for auxiliary supervision. For the online track, we introduce SwiftNet-Lookahead, which inserts a single bounded-lookahead module before a strictly causal iterative separator and keeps the total system latency at 96 ms. For the offline track, we use a frame-level enrollment cross-attention USEF-TFGridNet with a magnitude-domain fusion stage that trades off perceptual quality and speaker fidelity. In the official evaluation, SwiftNet-Lookahead ranks second in Track~1 and USEF-TFGridNet ranks fifth in Track~2, both exceeding the challenge baselines. These results suggest that real-data-oriented training and track-specific modeling are effective for conversational TSE.