Search papers, labs, and topics across Lattice.
Zhejiang University
9
0
9
9
CSAVocoder achieves real-time spatial audio generation with enhanced fidelity by effectively integrating dynamic spatial cues, outperforming traditional vocoders.
Current MLLMs falter in integrating multiple basketball knowledge domains, while BasketballSkills showcases superior performance through structured skill composition.
Robots can mimic human actions but fail to grasp the underlying intent, with performance collapsing when faced with novel tasks that require true understanding.
Reinforcement learning strategies can enable legitimate receivers to achieve near-optimal secrecy rates in competitive RIS auctions, significantly outpacing conventional bidding approaches.
Learning to avoid pseudo-robust features can drastically enhance model performance on unseen adversarial examples.
SwanTale achieves unprecedented expressiveness in multi-speaker audio generation, outperforming existing models in both instruct and zero-shot tasks.
SwanVoice leaps ahead in zero-shot TTS by nailing expressive, multi-speaker dialogue with a single model, finally bridging the gap between monologue quality and conversational coherence.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.