Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong
2
0
3
Zero-shot adaptation can reduce word error rates by over 0.6% while achieving nearly 10 times faster real-time processing for elderly speech recognition.
By explicitly modeling 3D space with learned spatial audio representations, JAEGER enables AV-LLMs to perform joint spatial grounding and reasoning far beyond the capabilities of 2D-centric models.