Search papers, labs, and topics across Lattice.
4
0
6
29
Surprisingly, simple "training-free" methods for combining vision and language beat more complex approaches for satellite image search, offering a scalable baseline for Earth observation.
MLLMs, without any training, can beat specialized models at image retrieval, especially when the target domain differs from the training data of those specialized models.
Achieve state-of-the-art open-vocabulary segmentation by distilling a sliding-window ViT into a single-pass architecture, enabling efficient high-resolution reasoning.
Just a handful of annotated examples can dramatically close the performance gap between zero-shot and fully supervised open-vocabulary segmentation.