Tsinghua AIMar 19, 2026arXiv:2603.19229

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

Huai-Zhou Jiang, Huaide Jiang, Yashdeep Chaudhary, Yash Chaudhary, Yuping Wang, Zehao Wang, Raghav Sharma, Manan Mehta, Yang Zhou, Lichao Sun, Lichao Sun, Zhiwen Fan, Zhengzhong Tu, Jiachen Li

AI Summary

The paper introduces NavTrust, a benchmark for evaluating the trustworthiness of embodied navigation agents under realistic corruptions in RGB, depth, and language instructions. Seven state-of-the-art VLN and OGN models were evaluated, revealing significant performance drops when exposed to these corruptions. The authors also tested four mitigation strategies and deployed the models on a real robot, demonstrating improved robustness.

Key Contribution

Embodied navigation agents, already struggling, fall apart when faced with the kinds of messy, real-world sensor and instruction corruptions that NavTrust now exposes.

Abstract

There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigation (OGN), where agents navigate to a specified target object. However, existing work primarily evaluates model performance under nominal conditions, overlooking the potential corruptions that arise in real-world settings. To address this gap, we present NavTrust, a unified benchmark that systematically corrupts input modalities, including RGB, depth, and instructions, in realistic scenarios and evaluates their impact on navigation performance. To our best knowledge, NavTrust is the first benchmark that exposes embodied navigation agents to diverse RGB-Depth corruptions and instruction variations in a unified framework. Our extensive evaluation of seven state-of-the-art approaches reveals substantial performance degradation under realistic corruptions, which highlights critical robustness gaps and provides a roadmap toward more trustworthy embodied navigation systems. Furthermore, we systematically evaluate four distinct mitigation strategies to enhance robustness against RGB-Depth and instructions corruptions. Our base models include Uni-NaVid and ETPNav. We deployed them on a real mobile robot and observed improved robustness to corruptions. The project website is: https://navtrust.github.io.

Eval Frameworks & Benchmarks Multimodal Models Robotics & Embodied AI

Citation Metrics

Citations0

Influential citations0

References39

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

Related Papers