Search papers, labs, and topics across Lattice.
This study investigates the robustness of cross-modal perception in underwater robots by integrating pretrained visual foundation models with sonar data under varying levels of visual degradation. The researchers employed a controlled benchmark to assess the performance of different fusion strategies, revealing that degradation-aware gated fusion significantly enhances detection accuracy in extreme conditions. Key findings demonstrate a 33.5% relative improvement in balanced accuracy when adapting the fusion mechanism to the reliability of each modality, highlighting the importance of dynamic modality contribution in challenging underwater environments.
Under extreme visual degradation, adaptive fusion of sonar and visual data boosts underwater detection accuracy by over 33%, revealing the critical role of modality reliability.
Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary information that is less affected by optical visibility, prior visual-sonar research has largely focused on feature alignment and nominal detection performance. We investigate cross-modal robustness as visual reliability deteriorates and assess whether pretrained visual foundation-model representations can be complemented by sonar under severe degradation. We use frozen DINOv2 as the visual encoder and construct a controlled five-level benchmark ranging from clean to extreme visual conditions. We compare conventional visual detection, frozen foundation-model representations, sonar context, fixed multimodal fusion, clean-trained adaptive gating, and degradation-aware gated fusion. Our method trains the fusion mechanism across the full range of degradation while keeping the visual and sonar encoders frozen, allowing modality contributions to adapt without fine-tuning the pretrained backbone. Under extreme combined degradation, the DINOv2 baseline achieves 0.4610 balanced accuracy, while degradation-aware visual-sonar fusion reaches 0.6152, a 33.5% relative improvement. The learned sonar contribution increases from 14.2% under clean conditions to 41.3% under extreme degradation, demonstrating adaptive redistribution of cross-modal reliance. Fusion provides the largest gains under severe turbidity and blur, whereas color attenuation alone yields little additional benefit. These results show that foundation-model representations remain valuable but insufficient under severe information loss, and that explicitly adapting fusion to modality reliability can improve robust underwater multimodal perception.