Search papers, labs, and topics across Lattice.
This paper introduces SAGE, a novel EEG-guided soft gating framework designed to enhance target speaker extraction during in-trial auditory attention switching. By employing a switch-aware gating module that dynamically selects and fuses candidate speech streams, SAGE effectively mitigates the challenges posed by neural noise and latency discrepancies. The method achieves significant performance improvements, with an 8.67 dB SI-SDR and 88.24% STOI, while reducing average switching latency to just 2.04 seconds, demonstrating its robustness in dynamic auditory environments.
SAGE reduces switching latency to 2.04 seconds while achieving superior target speaker extraction performance in the presence of neural noise and attention switching.
EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.