Search papers, labs, and topics across Lattice.
7
0
5
0
Domain shifts can cripple RMOT performance, revealing that the real challenge lies in maintaining stable associations between language and visual tracking rather than just detecting objects.
FUSION's innovative approach to integrating multi-cue features allows for robust tracking across diverse viewpoints, setting a new standard in pedestrian association and tracking.
Timage transforms the way we align text and images, achieving superior multimodal reasoning with a modest model size that outperforms larger competitors.
ExDet achieves state-of-the-art performance in open-domain open-vocabulary detection while significantly reducing training costs through innovative cross-modal techniques.
Real-time object detectors can achieve cross-domain generalization without any extra inference overhead by leveraging collaborative evidence modeling during training.
Object detection gets a flexible upgrade: now you can specify objects with text *and* images, opening the door to more intuitive and practical real-world applications.
Freezing your visual encoder and carefully nudging the text embeddings lets you continually teach an object detector new tricks without catastrophic forgetting.