Search papers, labs, and topics across Lattice.
This paper outlines the ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge, which evaluates the effectiveness of audio-visual speech enhancement (AVSE) in two distinct tracks: one with real-world overlapping speech and another with synthetic audio mixtures. The challenge addresses the limitations of existing protocols that rely on clean references and ideal conditions, providing a more realistic assessment of AVSE performance in naturalistic settings. Key results reveal that the baseline model achieved an SI-SDR of -4.069 dB and an STOI of 0.388 for real-world mixtures, highlighting the challenges of enhancing speech under overlapping conditions.
Real-world audio-visual speech enhancement struggles, with baseline models achieving only -4.069 dB SI-SDR in challenging overlapping scenarios.
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge evaluates two related settings. Track~1 comprises two scenarios: real-world mixtures recorded with two speakers speaking simultaneously, without a corresponding clean reference signal, and synthetic remixes obtained by manually mixing the separately recorded speech of two speakers, with a clean reference signal available; Track~2 reuses audio but pairs it with a degraded target video and contains additional 3-m far-field recordings. The speakers in the development and test sets are disjoint. Evaluation metrics include clean-waveform fidelity, learned quality estimates, transcription accuracy, and speaker identification. In the remix task on the development set, the baseline model achieved an SI-SDR of $-4.069$~dB and an STOI of $0.388$ on Track~1, and an SI-SDR of $-2.851$~dB and an STOI of $0.470$ on Track~2. We release the AV-ConvTasNet checkpoints, the offline evaluator, and the official baseline results on the development and test sets.