Search papers, labs, and topics across Lattice.
Karlsruhe Institute of Technology
4
0
6
High speech overlap isn't the primary challenge in cocktail-party scenarios; innovative audio-visual strategies and large language models can cut recognition errors by 57%.
Disfluencies are not just noise; they carry crucial meaning that, when ignored, significantly degrades translation quality.
Domain-adapted SpeechLLMs can be tricked into revealing sensitive information by transcribing phonetically similar words from their context or training data, even when a different word is spoken.
Current speech translation evaluation metrics are blind to critical speech-specific information, even when given the audio signal.