Search papers, labs, and topics across Lattice.
This study investigates the utilization of spatial information in vision-language models (VLMs) by employing direction patching, a causal intervention method. The findings reveal that the influence of spatial encodings on answer logits primarily manifests in mid-to-deep layers, with text chain-of-thought inhibiting immediate transport of object-word associations, while visually grounded prompts facilitate it. The research highlights that the encoding-grounding gap can be reframed as an issue of conditional transport, providing insights into how VLMs process spatial information differently across various contexts.
Spatial encodings in vision-language models only influence answers at deeper layers, revealing a nuanced transport mechanism that challenges existing assumptions about their functionality.
Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this with direction patching, a class-conditioned causal intervention applied across layers, token positions, and prompt formats. Using spatial-ID directions constructed following prior encoding evidence, we find that causal influence on answer logits emerges only at mid-to-deep depths. Text chain-of-thought suppresses immediate object-word argmax-level transport in most models, while visually grounded prompts keep it open. Positive target-logit gain can remain below the argmax threshold, and transport can re-emerge at the final prefix token or at the answer step in deeper layers. Across the ten VLMs we study, these local effects form descriptive transport patterns. Complementary experiments characterize how these patterns shift across datasets, attributes, and encoding amplitudes. Together, these results reframe the encoding-grounding gap as a problem of conditional transport in VLMs.