Search papers, labs, and topics across Lattice.
6
0
7
14
Achieving a 57.6% success rate on RoboCasa365, Xiaomi-Robotics-1 sets a new standard for vision-language-action models in real-world robotic manipulation.
MLLMs falter in fine-grained interpersonal reasoning, but integrating visual cues and social roles can dramatically boost their performance.
MLLMs struggle to juggle proactive tasks and reactive queries in dynamic video streams, but a simple agentic framework can significantly improve their coordination without any training.
Ditch the clunky architectures: a single diffusion model can now handle vision, language, and robot control to achieve SOTA manipulation performance.
A practical VLA model, LLaVA-VLA, achieves strong generalization and versatility on a new benchmark, CEBench, while running on consumer-grade GPUs, eliminating the need for costly pre-training.
By aligning latent representations with multiple visual foundation models, FRAPPE offers a more scalable and data-efficient way to imbue generalist robotic policies with robust world-awareness.