Search papers, labs, and topics across Lattice.
Chinese University of Hong Kong (Shenzhen), China
6
0
8
4
Real-world conversational dynamics significantly challenge target speaker extraction, revealing that even advanced systems struggle with natural overlap and noise.
By focusing on low/mid-frequency patterns, UniSkip-Mamba outperforms existing models in forgery detection, achieving a remarkable 63.4% AP@0.95 on LAV-DF.
MG-RWKV achieves state-of-the-art TFL performance with a groundbreaking O(T) complexity, redefining efficiency in audio-visual content authenticity verification.
A tool-augmented agent achieves high compilation rates but reveals a shocking 29-point gap in semantic faithfulness, challenging the reliability of existing evaluation metrics.
FlowTrain redefines VLM training efficiency, achieving up to 1.7x throughput improvements by decoupling execution and optimizing resource allocation.
CueNet achieves robust audio-visual speaker extraction under visual degradation by cleverly disentangling and integrating speaker information, acoustic synchronisation, and semantic synchronisation cues, without needing training on degraded visual data.