Search papers, labs, and topics across Lattice.
This paper introduces BiTSE, a binaural target speaker extraction framework designed to isolate desired speech in noisy multi-talker environments, crucial for augmented reality microphone arrays. By utilizing direction-of-arrival cues and voice activity information, BiTSE employs a DoA-aware attention mechanism, timestamp-based masking, and a two-stage loss optimization to significantly enhance signal fidelity. Evaluations on the SPEAR challenge dataset reveal that BiTSE outperforms traditional methods, achieving superior perceptual quality in speech extraction.
BiTSE achieves unprecedented improvements in speech extraction fidelity by effectively leveraging spatial and temporal cues in noisy environments.
Isolating a desired speech signal in noisy multi-talker conversational scenarios is a key requirement for augmented reality (AR) wearable microphone array systems. In this work, a binaural target speaker extraction (TSE) framework, termed BiTSE, is proposed. It leverages both spatial and temporal cues, specifically the direction-of-arrival (DoA) of the target speaker and corresponding voice activity information, to guide the extraction process. Built upon a binaural signal denoising architecture, our model integrates three key enhancements: (i) a DoA-aware attention mechanism using cyclic positional embeddings, (ii) a timestamp-based masking strategy that utilizes speaker activity to suppress non-target segments, and (iii) a novel two-stage loss optimization strategy that first trains the model for robust denoising and then fine-tunes it to improve perceptual quality. Evaluations on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset demonstrate that the proposed BiTSE consistently improves upon conventional approaches, leading to enhanced signal fidelity and perceptual quality.