Search papers, labs, and topics across Lattice.
The REAL-TSE Challenge at IEEE SLT 2026 focuses on advancing target speaker extraction (TSE) from real-world conversational recordings, emphasizing the complexities of natural speech such as overlap, noise, and channel mismatch. This challenge features two tracks鈥擮nline for low-latency extraction and Offline for full-context processing鈥攁llowing for a comprehensive evaluation of systems under varied conditions. Key findings reveal the effectiveness of different approaches through metrics like Token Error Rate and Speaker Similarity, providing valuable insights for future TSE benchmarks.
Real-world conversational dynamics significantly challenge target speaker extraction, revealing that even advanced systems struggle with natural overlap and noise.
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.