Search papers, labs, and topics across Lattice.
This paper introduces a unified framework for Audio-Visual Multi-Modal Learning (AVMML) that leverages a task-disentangled Low-Rank Adaptation (LoRA) mechanism to enhance multi-task collaboration. By integrating task-specific and shared knowledge through a combination of low-rank matrices and modulation strategies, the framework effectively mitigates mutual interference observed in naive joint training approaches. The results demonstrate that this method not only outperforms existing unified AVMML models but also surpasses many task-specific models on select tasks, highlighting its versatility and effectiveness.
Task-disentangled LoRA enables seamless integration of audio-visual tasks, outperforming both unified and task-specific models in multi-modal learning.
Inspired by human multi-modal perception, Audio-Visual Multi-Modal Learning (AVMML) integrates auditory and visual information to leverage complementary cross-modal cues, enabling more robust and comprehensive scene perception. Existing studies predominantly tackle each AVMML task in isolation, which stands in stark contrast to humans'unified cognitive capacity for handling versatile perception. However, naive joint training across multiple AVMML tasks often suffers from mutual interference, arising from the intricate inter-task relationships. To address this, we propose a unified framework that simultaneously accommodates versatile AVMML tasks. Specifically, benefiting from powerful representation and generalization capabilities of large language models, we design a task-disentangled Low-Rank Adaptation (LoRA) mechanism that enables dynamic integration of both task-specific and task-shared knowledge, thereby facilitating effective multi-task collaboration. The proposed task-disentangled LoRA comprises three components: a task-general low-rank matrix, task-specific modulation matrices, and cross-task collaboration experts, which respectively capture universal audio-visual knowledge, decouple task-specific pattern, and exploit inherent inter-task correlations. By unifying explicit collaboration from both model and task perspectives, our approach not only surpasses existing unified audio-visual models across multiple AVMML tasks, but also outperforms most task-specific models on certain AVMML tasks.