Search papers, labs, and topics across Lattice.
2
0
3
1
Fine-grained spatio-temporal reasoning in audio-language models can dramatically enhance the understanding of complex multi-event audio sequences.
Post-training with LoRA can boost accuracy in some models while hindering others, revealing the nuanced interplay between architecture and adaptation in audio-dependent tasks.