Search papers, labs, and topics across Lattice.
This paper introduces LiDAR-SAM2, a novel framework that leverages a 2D video foundation model to automatically generate high-quality, temporally consistent labels for 4D LiDAR segmentation without human intervention. By employing multi-view projection and spatio-temporal aggregation, the framework produces semantic and panoptic labels that closely match the quality of fully annotated datasets, even with minimal input. The results indicate that models trained on these automatically generated labels can achieve performance levels comparable to those trained on complete ground-truth supervision, significantly alleviating the annotation burden in 3D and 4D scene understanding tasks.
Automating LiDAR annotation can cut down the labor-intensive process of data labeling while achieving near-human accuracy in segmentation tasks.
Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense annotation. This motivates a simple question of whether high-quality LiDAR training data can be produced automatically, without any human labeling. To this end, we introduce LiDAR-SAM2, a framework that turns a 2D video foundation model, SAM2, into a scalable source of supervision for the 4D LiDAR domain. On the data side, it automatically generates temporally coherent LiDAR-level labels from SAM2 video masks through multi-view projection and spatio-temporal aggregation. On the modeling side, a tailored modality interface and a two-stage learning objective adapt SAM2's video segmentation kernel to spatio-temporal LiDAR structure, so that a single click per object yields a consistent mask track across the sequence. Trained with no human LiDAR annotation, LiDAR-SAM2 produces semantic and panoptic labels on SemanticKITTI that approach the quality of full human annotation from only a few points, and models trained on these labels approach the performance of full ground-truth supervision. This positions LiDAR-SAM2 as a scalable labeling tool that substantially reduces the annotation burden for 3D and 4D scene understanding.