Search papers, labs, and topics across Lattice.
Affiliation:
7
0
7
Visual tracks can double the success rates of robot tasks by providing a powerful interface between control and visual prediction.
Revealing robot motion in video models can transform how we predict and control robotic actions, achieving high fidelity with minimal training data.
Scaling visuomotor context to 8K timesteps enables robots to master complex tasks and adapt in real-time, outperforming previous models by a staggering margin.
Policies trained in SimFoundry's automated environments achieve up to 40% higher success rates in real-world tasks by leveraging affordance-preserving scene variations.
VLMs can learn to actively reason and plan in 3D environments by distilling view graphs from self-exploration trajectories, enabling them to surpass even larger models like GPT-4 Pro and Gemini 1.5 Pro on interactive view planning.
LLM agents can appear to reason well (high entropy) while completely ignoring the input, and mutual information is a far better metric for catching this failure.
Robots can now learn from their mistakes in real-time via a novel reflective planning framework, leading to significant performance gains in long-horizon tasks.