Search papers, labs, and topics across Lattice.
This paper introduces CauseCollab, a causal unified and modality-agnostic network designed to enhance collaborative perception by addressing the challenges posed by heterogeneous sensor modalities and model architectures. By employing causal metric learning to disentangle semantic factors from modality-specific confounders, CauseCollab ensures cross-modal semantic consistency and minimizes error accumulation in diverse environments. Extensive experiments on the OPV2V and DAIR-V2X datasets reveal that CauseCollab outperforms existing methods, particularly in scenarios with substantial modality discrepancies.
Causal metric learning in CauseCollab dramatically improves semantic consistency across heterogeneous modalities, leading to state-of-the-art performance in collaborative perception tasks.
Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate this problem by mapping heterogeneous features into a shared protocol space; however, independently trained modality-specific converters often generate modality-specific pseudo-protocol distributions, leading to semantic inconsistency and error accumulation, which is particularly pronounced in scenarios with large modality discrepancies. To address this issue, we propose CauseCollab, a causal unified and modality-agnostic network. CauseCollab formulates representation learning in the protocol space from a causal perspective, explicitly disentangling semantic factors from modality-specific statistical confounders via causal metric learning. Meanwhile, CauseCollab adopts context-guided Unified Converter for heterogeneous modalities to ensure cross-modal semantic consistency. In addition, integrating new modalities only requires training adapters with minimal parameters. Extensive experiments on the OPV2V and DAIR-V2X datasets demonstrate that CauseCollab achieves state-of-the-art performance, with more significant gains in scenarios involving large modality gaps.