Search papers, labs, and topics across Lattice.
2
0
3
5
Experiments conducted on the UASpeech and TORGO corpora suggest that Conformer models trained using the DSI tokens outperform the comparable baseline HuBERT discrete/continuous features by statistically significant WER reductions.
By explicitly modeling 3D space with learned spatial audio representations, JAEGER enables AV-LLMs to perform joint spatial grounding and reasoning far beyond the capabilities of 2D-centric models.