Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong, Shenzhen
2
0
4
1
Fine-tuning large audio-language models for emotion recognition can be drastically improved by leveraging hyperbolic geometry, leading to better performance on class-imbalanced datasets.
Slash spoken dialogue system latency by up to 51% with a new architecture that lets the system "listen-while-thinking" and "speak-while-thinking."