Search papers, labs, and topics across Lattice.
2
0
4
8
GRA reveals that grounding VLA models in geometric representations from generated videos can outperform traditional methods that attempt to extract control signals from the same data.
Current multimodal agents are nowhere near human-level gaming, even when given precise semantic actions, revealing critical gaps in real-time interaction and context handling.