Search papers, labs, and topics across Lattice.
2
0
3
2
Query-grounded visual sampling can boost long video understanding in LVLMs by over 11% without the need for extensive training.
Revitalizing the latent edge-sensitivity of DINO with SAM's structural priors yields state-of-the-art open-vocabulary segmentation, especially in cluttered scenes.