Search papers, labs, and topics across Lattice.
This paper introduces QIANGDA, a benchmark for audio-visual target speaker extraction (AV-TSE) that includes 77 scenes and over 7,500 clips, designed to improve the accuracy of speaker identification by leveraging synchronized multi-view video. The authors also curate VOXBLINK2-AVSE, a large dataset of audio-lip pairs, to enhance training and evaluation of their extraction model. The resulting extractor achieves a competitive character error rate (CER) of 0.2261 and demonstrates significant improvements in output correctness and success rates, highlighting the effectiveness of their approach in utilizing visual cues for speaker extraction.
Achieving over 82% output correctness, this new benchmark and model redefine the standards for audio-visual target speaker extraction by effectively integrating visual cues.
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. We introduce QIANGDA, a Mandarin AV-TSE benchmark of jointly recorded real two-speaker mixtures with synchronized multi-view video. Each scene also contains preceding A-only and B-only stages that provide in-scene speaker references. It contains 77 scenes and 7,598 clips (11.84 hours), including 6,042 dual-annotated mixtures. After processing there leave 6,038 evaluable mixtures and 12,076 target-speaker rows. We additionally curate VOXBLINK2-AVSE from VoxBlink2, comprising 250,828 synchronized audio--lip-ROI pairs from 28,421 identities and 766.17 hours of speech. Our extractor uses frozen, 1,280-dimensional projected AV-HuBERT features, target-conditioned training, and layer-wise feature modulation. We jointly evaluate content with Qwen3-ASR-1.7B CER and target identity with WeSpeaker ResNet34 plus Overlapped Speech Detection (OSD). On the complete manifest, the best archived checkpoint obtains 0.2261 CER, 82.22% strict output correctness, and 69.53% both-output strict success.