Search papers, labs, and topics across Lattice.
This study employs causal tracing to identify the visual information that influences the outputs of vision-language models (VLMs), revealing that significant causal tokens often originate from outside the expected target regions. The analysis across various models and corruption scenarios indicates that high performance in multimodal tasks does not equate to spatially localized causal representations. Additionally, the research uncovers that while VLMs rely on visual cues to interpret structures, they struggle to maintain visual coherence when these cues are removed, highlighting a disconnect between perception and reasoning in these models.
Causal tracing reveals that crucial visual tokens in VLMs often come from unexpected areas, challenging assumptions about spatial localization in multimodal reasoning.
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.