Search papers, labs, and topics across Lattice.
This paper introduces MoSaiC, a novel framework for self-supervised point cloud video representation learning that enhances 3D dynamic scene understanding. By integrating Curriculum Motion-Saliency Masking, Normal-Flow Motion modeling, and Cross-view Token Consistency Prediction, MoSaiC effectively captures both appearance and motion dynamics in point cloud videos. Extensive evaluations across various tasks, including action recognition and semantic segmentation, reveal that MoSaiC significantly outperforms existing methods, highlighting its potential in advancing point cloud video analysis.
MoSaiC captures motion dynamics in point cloud videos more effectively than existing methods, leading to superior performance in action recognition and segmentation tasks.
Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud video representation learning. MoSaiC couples three components: Curriculum Motion-Saliency Masking (CMSM), which guides the masking process toward motion-salient tokens under a curriculum schedule; Normal-Flow Motion (NFM) modeling, which supervises the local rigid rotation of each token in the Lie algebra so(3) as an explicit geometric motion target; and Cross-view Token Consistency Prediction (CTCP), which enforces consistency between two complementary masked views at the token level. Together, these components allow MoSaiC to effectively capture both appearance and motion dynamics. Extensive experiments on multiple downstream tasks, including action recognition, temporal action segmentation, and point-level semantic segmentation, demonstrate the effectiveness of our approach.