Search papers, labs, and topics across Lattice.
This study investigates the relational reasoning capabilities of Vision-Language Models (VLMs) by analyzing the Qwen3-VL-4B model's encoding of visual information across different depths. Using a synthetic dataset of geometric shapes, the researchers designed specific queries to assess the models' reliance on language cues versus visual evidence. The findings reveal that while VLMs exhibit some genuine visual reasoning, they predominantly rely on language shortcuts, highlighting a significant gap in their understanding of visual relations.
VLMs blend real visual reasoning with a heavy reliance on language shortcuts, raising questions about their true understanding of visual relations.
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.