Search papers, labs, and topics across Lattice.
2
0
4
4
Decoupling perception from reasoning in visual tasks leads to a remarkable 93.2% accuracy on V-Star, showcasing a new paradigm for fine-grained visual reasoning.
Multimodal models can "see" the image but still fail at reasoning because the visual input distracts the routing mechanism from activating the right experts.