Search papers, labs, and topics across Lattice.
3
0
6
4
Decoupling perception from reasoning in visual tasks leads to a remarkable 93.2% accuracy on V-Star, showcasing a new paradigm for fine-grained visual reasoning.
Multimodal models can "see" the image but still fail at reasoning because the visual input distracts the routing mechanism from activating the right experts.
Turns out, MLLMs struggle with manufacturing tasks not because they can't "see," but because they lack the domain-specific knowledge to understand what they're looking at.