Search papers, labs, and topics across Lattice.
4
0
4
Bridging the gap between domains with tailored noise and selective inpainting can revolutionize few-shot object detection performance.
Counterintuitively, pushing *away* certain image tokens from their corresponding text embeddings boosts CLIP's few-shot cross-domain performance.
Fine-tuning vision-language models for cross-domain few-shot learning makes them *worse* at distinguishing between classes due to an exacerbated "attention sink" problem, but a simple token re-weighting scheme can fix it.
Object detectors in new visual domains suffer from "astigmatism," but mimicking the human eye's foveal vision can bring them into focus.