Search papers, labs, and topics across Lattice.
2
0
3
5
Different views of the same problem reveal hidden reasoning paths, enabling VLMs to achieve unprecedented accuracy in multimodal reasoning tasks.
Current multimodal large language models struggle with active visual observation, achieving less than 11% accuracy on a benchmark designed to measure this critical capability.