Search papers, labs, and topics across Lattice.
NTT, Inc., Japan
7
0
6
3
Real-world conversational dynamics significantly challenge target speaker extraction, revealing that even advanced systems struggle with natural overlap and noise.
SphereVBx achieves superior clustering accuracy in speaker diarization while significantly simplifying the process, making it a game-changer for EEND-VC applications.
NAR-MBR decoding achieves superior speech recognition accuracy while being faster than autoregressive methods, redefining efficiency in real-time applications.
LLMs can outperform humans in predicting the next speaker in meetings, even without audio or visual data.
Neural networks trained with C2D-projected data achieve superior performance in real-world speech enhancement, outperforming state-of-the-art methods by effectively bridging the gap between close and distant microphone recordings.
Achieving 70% of the benefits of ideal tight-label training, this method transforms how speaker diarization models handle loose annotations.
Real-world audio scene understanding takes a leap forward with a benchmark that tackles the complexities of overlapping sound events and absent targets.