Search papers, labs, and topics across Lattice.
2
0
4
0
Domain shifts can cripple RMOT performance, revealing that the real challenge lies in maintaining stable associations between language and visual tracking rather than just detecting objects.
Injecting bounding box information directly into the visual modality of video LLMs slashes token costs by up to 93% and unlocks better spatial reasoning.