Search papers, labs, and topics across Lattice.
4
0
4
2
The challenge provides a benchmark dataset, pretrained baseline models, and an evaluation framework to advance face--voice association, highlighting the need to foster the development of models that capture identity-specific aspects beyond language and gender.
Real-world speaker identification can thrive even with missing modalities and multilingual contexts, challenging the status quo of multimodal systems.
Top systems in the ESDD2 challenge achieved a staggering Macro-F1 score of 0.8775, revealing the power of modular design and self-supervised learning in audio deepfake detection.
Environmental sound deepfakes are a rising threat, and this challenge reveals the current state-of-the-art in detecting them, highlighting both the progress and remaining gaps.