Search papers, labs, and topics across Lattice.
Affiliation:
6
0
7
4
Automating clinical note generation from speech could drastically reduce healthcare workers' documentation time while preserving essential patient information.
High speech overlap isn't the primary challenge in cocktail-party scenarios; innovative audio-visual strategies and large language models can cut recognition errors by 57%.
Disfluencies are not just noise; they carry crucial meaning that, when ignored, significantly degrades translation quality.
Domain-adapted SpeechLLMs can be tricked into revealing sensitive information by transcribing phonetically similar words from their context or training data, even when a different word is spoken.
Current speech translation evaluation metrics are blind to critical speech-specific information, even when given the audio signal.
Text prompts might be inflating your SLLM's performance: spoken prompts reveal a significant performance gap, especially in low-resource languages.