Search papers, labs, and topics across Lattice.
This paper introduces a robust framework for dynamic object segmentation that integrates multimodal cues, including 2D point tracks, 3D reconstruction, and semantic information, to produce precise and complete dynamic masks. By employing a network that combines Transformer architectures with feature clustering aggregation modules, the model adaptively classifies static and dynamic features while reducing the impact of degradation. Extensive experiments reveal that this approach achieves state-of-the-art performance in both dynamic object segmentation and static scene reconstruction tasks, addressing critical limitations of existing methods.
Achieving state-of-the-art dynamic object segmentation by seamlessly integrating multimodal cues could redefine how we approach visual scene understanding.
Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive to reconstruction errors. To address these limitations, we present a dynamic object segmentation framework that can generate both precise and complete dynamic masks by integrating multimodal cues including 2D point tracks, 3D reconstruction, and semantic information. We design a network combining Transformer architectures with feature clustering aggregation modules to perform static/dynamic classification of multimodal feature trajectories. It enables the model to adaptively determine which type of feature should dominate based on the characteristics of each scene, while also mitigating the impact of feature degradation. Additionally, we introduce a novel point-query-based SAM post-processing method capable of handling multiple objects within a single mask. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in both dynamic object segmentation and static scene reconstruction tasks.