Search papers, labs, and topics across Lattice.
8
0
11
8
Human demonstrations can yield over 10x the recovery data for robots, dramatically enhancing their ability to recover from failures in real-world tasks.
Achieving 89.4% success in single-view embodied visual tracking, ReferTrack outperforms multi-camera systems by grounding tracking decisions in real-time visual data.
Hy-Embodied-VLM-1.0 outperforms its predecessor by 8.4% while activating only a fraction of the parameters, redefining efficiency in embodied agents.
Generating visually faithful driving simulations just got a boost with a novel framework that stabilizes error accumulation and enhances realism in closed-loop scenarios.
Transforming 2D visual features into interpretable 3D representations unlocks a new level of spatial intelligence in MLLMs, offering unprecedented insights into their internal workings.
Twitter strips C2PA provenance data from AI-generated images, making it impossible to cryptographically verify their origin on the platform.
GPT-Image-2 can so seamlessly forge documents that neither humans nor the model itself can reliably tell the difference.
Forget boring ads: this new method uses creative knowledge to generate videos that actually match product features and move realistically.