Search papers, labs, and topics across Lattice.
This paper introduces DRGBT-1K, a large-scale benchmark specifically designed for Dynamic RGBT tracking, which captures 1,045 sequences in real-world scenarios with diverse modalities and viewpoints. The benchmark includes comprehensive annotations for fine-grained evaluation across 24 target categories and 15 challenge attributes, addressing the limitations of existing benchmarks that fail to assess tracker robustness under dynamic conditions. By evaluating 20 multimodal tracking methods under a unified protocol, the authors provide a critical resource for advancing research in unaligned multimodal tracking and UAV-ground collaborative tracking.
A groundbreaking benchmark that captures the complexities of real-world dynamic tracking with 795K RGBT frame pairs and extensive annotations for robust evaluation.
Dynamic RGBT (DRGBT) tracking aims to continuously localize a target when the available sensing modalities and observation platforms vary over time. Compared with conventional RGBT tracking with fixed RGBT inputs and a fixed observation platform, DRGBT tracking is more consistent with real-world collaborative perception systems, where targets may be observed by heterogeneous sensors from different viewpoints. However, existing benchmarks are still insufficient for systematically evaluating tracker robustness under real dynamic modality variations and cross-platform transitions. To address this limitation, we make the following contributions. 1) We construct DRGBT-1K, a large-scale high-quality benchmark for DRGBT tracking. It contains 1,045 sequences captured entirely in real-world scenarios and 795K RGBT frame pairs collected using UAVs and handheld RGBT devices, encompassing diverse real-world scenes, pronounced viewpoint changes, modality variations, and target appearance discontinuities. 2) We provide comprehensive annotations for fine-grained evaluation, including dense bounding boxes, target category labels, challenge attributes, frame-level modality labels and platform labels. DRGBT-1K covers 24 target categories, more than 15 scene types and 15 challenge attributes. 3) We establish a comprehensive benchmark by evaluating 20 representative multimodal tracking methods, including conventional RGBT trackers, modality-missing RGBT trackers, and DRGBT trackers under a unified evaluation protocol. 4) We release an unaligned version of DRGBT-1K and derive UGVT-1K to support broader research on unaligned multimodal tracking and UAV-ground collaborative tracking. 5) We develop an online evaluation platform for DRGBT-1K and provide a leaderboard that collects all methods evaluated on this benchmark.