Search papers, labs, and topics across Lattice.
This paper introduces SkyVLaM, a multimodal large language model designed to enhance UAV video understanding in remote sensing by addressing the challenges of small, ambiguous targets and dynamic perspectives. The model innovatively constructs sparse tokens from patch-level video representations and employs a temporal basis perceiver to facilitate complementary temporal cues, leading to improved query-conditioned segmentation. Experimental results demonstrate that SkyVLaM significantly optimizes visual token allocation and enhances segmentation performance across various UAV video scenarios.
SkyVLaM revolutionizes UAV video understanding by effectively balancing sparse and dense token processing, leading to superior segmentation of visually ambiguous targets.
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.