Search papers, labs, and topics across Lattice.
This paper introduces an embodied multimodal grounding framework that enhances open-vocabulary mobile manipulation by integrating active multi-view Semantic 3D Gaussian Splatting with a diffusion-based vision-language-action policy. The framework effectively aligns language, visual observations, and three-dimensional scene structure to facilitate few-shot manipulation in cluttered household environments. Experimental results demonstrate a significant improvement in long-horizon success rates, achieving 60% compared to 40% for existing methods, highlighting the framework's robustness against various challenges such as occlusion and viewpoint variation.
Explicit 3D semantic grounding boosts manipulation success rates by 50% in cluttered environments, showcasing a leap in mobile manipulation capabilities.
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that integrates active multi-view Semantic 3D Gaussian Splatting (Semantic-3DGS), reachability-aware base positioning, and a diffusion-based vision-language-action policy. A task-driven local Semantic-3DGS serves as a shared interface across active sensing, language-conditioned 3D localization, obstacle-aware scene reasoning, base preparation, and semantic conditioning of the action model. To preserve pretrained action priors, the 3D semantic cues are injected only into the late action-expert blocks. In expanded 50-trial real-robot evaluations against representative vision-language-action (VLA) approaches, the full system achieves 60% long-horizon success compared with 40% for PointVLA and 28% for DexVLA, and reaches 74% success in heavily cluttered manipulation compared with 52% for the single-view variant and 46% for PointVLA. It also maintains 75% success under a 75 cm height shift and eliminates photo-induced false grasps. These results indicate that explicit, refreshable 3D semantic grounding can improve robustness under clutter, occlusion, viewpoint variation, and embodiment constraints.