Search papers, labs, and topics across Lattice.
Fondazione Bruno Kessler
4
0
5
16
Fixed 30-second segmentation emerges as the key to robust long-form speech instruction following, outperforming other methods.
Vulnerabilities in speech models are not just a problem for English; they worsen in other languages and with spoken inputs, revealing a critical oversight in AI safety.
Skip the training: SimulU achieves state-of-the-art simultaneous speech translation by cleverly exploiting pre-trained models, opening the door to truly plug-and-play multilingual communication.
Text prompts might be inflating your SLLM's performance: spoken prompts reveal a significant performance gap, especially in low-resource languages.