Search papers, labs, and topics across Lattice.
Institute of Computing Technology, University of Chinese Academy of Sciences
2
0
3
Explicitly grounding evidence in spatial relation tasks can boost VLM performance by nearly 12 points, transforming how we approach visual reasoning.
A staggering 12.73-26.25% of correct decisions in multimodal spatial reasoning are made without proper credit to the supporting images, exposing flaws in current evaluation methods.