Search papers, labs, and topics across Lattice.
4
0
4
27
This work presents a multimodal conversational-context framework for LLM-based ASR that integrates a scenario-controlled data pipeline, scalable multimodal context training, and systematic evaluation and introduces MM-ContextASR Bench, which evaluates contextual understanding and entity error correction across five scenarios.
This technical retrospective offers a different account of the CosyVoice lineage, from CosyVoice through CosyVoice 2 and CosyVoice 3 to Qwen-Audio-3.0-TTS, and formalizes this history through four interface dimensions---representation, ownership, availability, and gradient reach---and separate within-paper evidence from cross-paper comparison.
Correctly aligning audio descriptions with representations can boost classification performance by nearly 5 points, revealing the critical role of semantic refinement in audio tasks.
Despite advances in large audio language models, the best performer in temporal audio grounding only achieves a mere 31.2 mIoU, revealing critical limitations in current capabilities.