Search papers, labs, and topics across Lattice.
Wadhwani AI, IIIT Hyderabad
3
0
5
25
Naive translation of audio descriptions fails to capture cultural nuances, revealing a critical gap in accessibility for India's Blind and Low Vision communities.
Attention dynamics reveal that MLLMs can fall back on language priors when visual context is disrupted, highlighting the fragility of multimodal integration.
Forget salient cues – now you can *steer* visual representations in ViTs with language, focusing on any object you want without hurting overall performance.