Search papers, labs, and topics across Lattice.
This paper introduces a scalable framework for learning urban navigation policies from in-the-wild egocentric videos, addressing the limitations of traditional data collection methods. By automatically annotating web videos with navigation semantics, the authors train a vision-language-action policy that enhances interpretability in navigation planning. Key results reveal a coherent long-tail structure in navigation performance, highlighting critical safety scenarios often overlooked in conventional datasets.
Uncovering a coherent long-tail structure in urban navigation reveals critical safety scenarios that traditional data approaches often miss.
Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learning point-goal urban navigation from web-scale in-the-wild egocentric videos while systematically exposing its long tail. The framework automatically annotates uncurated web videos with metric trajectories and structured navigation semantics, which are then used to train a vision-language-action policy for interpretable navigation planning. We characterize the long tail based on model performance and the distribution of perception-motion patterns, and employ reflection-based analysis to diagnose recurring failure modes. Experiments on web-video data and real-world urban navigation tasks demonstrate effective knowledge transfer from unconstrained videos and reveal coherent long-tail structures beyond aggregate navigation performance.