Search papers, labs, and topics across Lattice.
5
0
5
4
VLM-IE3D achieves state-of-the-art performance in 3D tasks by seamlessly integrating implicit and explicit geometric representations from RGB inputs.
Jointly optimizing visual token selection and LLM computation can drastically reduce inference costs while enhancing performance in multimodal tasks.
Video LLMs can significantly improve their QA performance by integrating spatio-temporal evidence, bridging the gap between accuracy and visual perception.
GeoProp achieves a remarkable 10.6% boost in real-world manipulation tasks by effectively grounding robot state in visual context, all while adding minimal complexity.
Robots can now autonomously adapt to camera changes without needing explicit calibration, significantly improving deployment flexibility.