Search papers, labs, and topics across Lattice.
2
0
4
7
Current VLA models can execute tasks based solely on visual cues, but they also risk following unauthorized cues, raising critical safety concerns.
Visual reasoning in multimodal models is highly variable, with accuracy plummeting over 10 points when feedback is corrupted, revealing hidden dependencies on visual states.