Search papers, labs, and topics across Lattice.
This paper introduces a structured evaluation protocol for LiDAR semantic segmentation that assesses deployment readiness across three critical dimensions: coarse-label evaluation, robustness to sensor corruptions, and domain generalization. The authors find that traditional fine-grained benchmarks fail to capture safety-relevant performance, with significant degradation observed under real-world conditions and architecture-dependent robustness characteristics. Their results highlight the inadequacies of current evaluation methods and propose a more practical framework for assessing LiDAR systems in autonomous vehicles and mobile robots.
Fine-grained benchmarks may mislead deployment readiness, as real-world performance is heavily impacted by label granularity and environmental corruptions.
LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despite the strong performance of recent state-of-the-art methods on standard benchmarks, existing evaluation protocols remain focused on clean, single-domain settings and fine-grained label taxonomies, leaving deployment readiness largely unassessed. Real-world systems must handle safety-critical label semantics, degraded sensing conditions, and cross-domain variability, yet no unified protocol currently addresses all three aspects together. In this paper, we propose a structured evaluation protocol that assesses the deployment readiness of LiDAR semantic segmentation models along three complementary dimensions: (i) coarse-label evaluation aligned with autonomous driving safety priorities, revealing how label granularity affects different methods; (ii) robustness under eight types of LiDAR corruptions designed to emulate real-world atmospheric, geometric, and sensor degradations; and (iii) domain generalization across datasets without adaptation. The evaluation includes inference speed measured on an embedded Jetson AGX Orin platform, directly reflecting deployment constraints. Our results show that fine-grained benchmark rankings do not always reflect safety-relevant performance, that all methods experience substantial degradation under corruptions with architecture-dependent robustness characteristics, and that current domain generalization remains insufficient for reliable deployment. These findings expose concrete gaps between benchmark performance and deployment readiness, and provide a reference protocol for more practically grounded evaluation of LiDAR semantic segmentation.