Search papers, labs, and topics across Lattice.
2
0
3
7
Tying intermediate reasoning directly to visual evidence allows VLMs to outperform larger models in spatial reasoning tasks.
OpenVLThinkerV2 leapfrogs existing open-source and proprietary multimodal models by using a novel Gaussian-based RL objective that ensures gradient equity across diverse visual tasks.