Search papers, labs, and topics across Lattice.
IIIT Delhi, India
3
0
5
2
Contextual entrainment in vision-language models reveals a duality that could fundamentally alter our understanding of multimodal interactions.
VLMs are often functionally blind, exploiting language priors instead of truly "seeing" visual data, and this problem paradoxically *worsens* as language models scale.
Existing citation recommendation benchmarks overestimate real-world performance because they fail to account for the temporal constraints of recommending citations for *new* papers.