Search papers, labs, and topics across Lattice.
This paper introduces CD-RMOT-Bench, a benchmark designed to evaluate Cross-Domain Referring Multi-Object Tracking (CD-RMOT), which assesses the ability of RMOT models to track objects based on natural-language expressions across different visual domains. The study reveals that existing models suffer significant performance degradation due to domain shifts, primarily due to unstable expression-conditioned temporal associations rather than object detection errors. The authors propose a Query-Centric Adaptation (QCA) framework that stabilizes the query space, providing a strong baseline for future research in robust language-guided tracking.
Domain shifts can cripple RMOT performance, revealing that the real challenge lies in maintaining stable associations between language and visual tracking rather than just detecting objects.
Referring multi-object tracking (RMOT) extends tracking from category-driven perception to language-guided understanding by grounding object trajectories in natural-language expressions. Despite recent progress, existing RMOT studies are largely conducted under in-domain settings, leaving the robustness of language-conditioned tracking under inevitable visual domain shifts unexplored. In this paper, we study Cross-Domain Referring Multi-Object Tracking (CD-RMOT), a new and challenging problem that evaluates whether an RMOT model trained on a labeled source domain can reliably follow natural-language expressions in an unlabeled target domain with different visual conditions. To support systematic study, we construct CD-RMOT-Bench, a unified benchmark that combines real clear-domain referring tracking data, aligned digital-twin variants, and real adverse-domain videos. CD-RMOT-Bench enables both controlled weather/viewpoint shift analysis and realistic synthetic-real transfer evaluation under a shared RMOT protocol. Further, we provide a Query-Centric Adaptation (QCA) framework, designed to stabilize the query space that bridges visual trajectories and referring expressions. Extensive experiments reveal that domain shifts severely degrade RMOT performance, where the failure is not merely caused by object detection errors but more critically by unstable expression-conditioned temporal association and target selection. QCA establishes a strong baseline, while CD-RMOT-Bench opens a new direction for robust language-guided tracking across visual domains.