Search papers, labs, and topics across Lattice.
Sun Yat-sen University
3
0
3
Injecting physics-inspired spatial priors into video MLLMs drastically enhances temporal stability and segmentation accuracy without sacrificing image-level performance.
COSMO redefines source-free domain adaptation by dynamically balancing the influence of source models and vision-language models, achieving unprecedented performance while mitigating source bias.
Achieve state-of-the-art person re-identification with only 20% of the data by explicitly teaching the model to "think" before matching identities.