Search papers, labs, and topics across Lattice.
This paper introduces a unified planning-learning framework for Unmanned Underwater Vehicles (UUVs) that enhances navigation in dynamic environments by integrating persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The framework employs a reinforcement learning policy for short-range tracking and reactive avoidance, while also learning a compact latent state representation from onboard sensor data to improve decision-making under partial observability. Experimental results show that the proposed method significantly outperforms traditional behavior tree and standard reinforcement learning baselines in terms of robustness and safety in complex underwater scenarios.
A novel hybrid planning-learning architecture enables UUVs to navigate dynamically changing underwater environments with unprecedented robustness and safety.
This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboard sonar and depth image observations, adapts a clearance-constrained global planner (GP) to provide long-horizon structure, and integrates a reinforcement learning (RL) policy to handle short-range tracking and reactive avoidance. To further support decision-making under partial observability, the system learns a compact latent state representation from onboard sensor data, encoding environmental structure, obstacle dynamics, and uncertainty. Behavior tree (BT) distillation with staged supervision is introduced to improve safety and training stability, while an uncertainty-calibrated distillation mechanism reweights teacher guidance using online latent-model uncertainty, emphasizing uncertain regimes during learning, with time-to-collision (TTC) and clearance cues remaining explicit in planning and local policy features. To demonstrate the efficacy of the framework, a reproducible multi-seed evaluation protocol is established in high-fidelity GPU-accelerated simulation using NVIDIA Isaac Sim, and performance is benchmarked against BT-only and standard RL baselines. The results obtained demonstrate improved robustness and safety under dynamic conditions, thus providing a general pipeline with a unified hybrid planning learning architecture and a reproducible methodology for robust UUV autonomy under partial observability.