Search papers, labs, and topics across Lattice.
This paper introduces the Multi-Skill Manipulation-Enhanced Mapping (MS-MEM) framework, which enhances scene understanding for service robots in cluttered environments by integrating uncertainty-aware mapping with active viewpoint selection, object pushing, and grasping. The approach employs a full-evidential grasp estimator that captures both grasp affordance and orientation uncertainty, allowing for a unified action selection pipeline that minimizes scene disturbance while improving mapping accuracy. Experimental results demonstrate that MS-MEM outperforms traditional single-skill and unconstrained methods, achieving superior mapping accuracy with significantly reduced scene disturbances.
MS-MEM achieves higher mapping accuracy while minimizing scene disturbances, showcasing the power of integrating multiple manipulation skills in robotic perception.
Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this paper, we propose Multi-Skill Manipulation-Enhanced Mapping (MS-MEM), an evidential framework for uncertainty-aware mapping that integrates active viewpoint selection, object pushing, and grasping. MS-MEM combines scene-level metric-semantic evidential belief estimators with an uncertainty-aware grasp representation. This representation is learned using a novel full-evidential grasp estimator that models both grasp affordance and orientation uncertainty. In our framework, candidate perception and manipulation actions are evaluated within a unified action selection pipeline using a common information gain criterion. For manipulation actions, we further introduce a collateral disturbance constraint (CDC) that discourages excessive changes to confident regions of the scene belief. This enables MS-MEM to select actions that effectively reduce map uncertainty while limiting collateral scene changes. Experimental results show that, compared with single-skill and unconstrained baselines that ignore scene disturbance, MS-MEM achieves higher mapping accuracy while substantially reducing scene disturbance, highlighting the synergistic effects of active viewpoint selection, push, and grasp actions.