Search papers, labs, and topics across Lattice.
The paper introduces HiMemVLN, a novel approach to improve zero-shot Vision-Language Navigation (VLN) using open-source LLMs by addressing the "Navigation Amnesia" problem. HiMemVLN incorporates a Hierarchical Memory System into a multimodal LLM to enhance visual perception recall and long-term localization. Experiments in simulated and real-world environments show that HiMemVLN nearly doubles the performance of existing open-source VLN methods.
Open-source VLN agents can nearly double their navigation success by remembering where they've been, thanks to a new hierarchical memory system.
LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs as navigators, which face challenges related to high token costs and potential data leakage risks. Recent efforts have attempted to address this by using open-source LLMs combined with a spatiotemporal CoT framework, but they still fall far short compared to closed-source models. In this work, we identify a critical issue, Navigation Amnesia, through a detailed analysis of the navigation process. This issue leads to navigation failures and amplifies the gap between open-source and closed-source methods. To address this, we propose HiMemVLN, which incorporates a Hierarchical Memory System into a multimodal large model to enhance visual perception recall and long-term localization, mitigating the amnesia issue and improving the agent's navigation performance. Extensive experiments in both simulated and real-world environments demonstrate that HiMemVLN achieves nearly twice the performance of the open-source state-of-the-art method. The code is available at https://github.com/lvkailin0118/HiMemVLN.