Search papers, labs, and topics across Lattice.
6
0
8
10
Counterintuitively, pushing *away* certain image tokens from their corresponding text embeddings boosts CLIP's few-shot cross-domain performance.
Fine-tuning vision-language models for cross-domain few-shot learning makes them *worse* at distinguishing between classes due to an exacerbated "attention sink" problem, but a simple token re-weighting scheme can fix it.
Agentic models can learn to trust their "gut" and rely less on external tools, leading to faster and more accurate reasoning.
Object detectors in new visual domains suffer from "astigmatism," but mimicking the human eye's foveal vision can bring them into focus.
CLIP struggles with fine-grained details in cross-domain few-shot learning, but a cycle-consistency method can fix its vision-language alignment and boost performance.
CLIP's "lost" text encoder layers actually contain valuable information for cross-domain few-shot learning, and a method to re-utilize them significantly boosts performance.