Search papers, labs, and topics across Lattice.
The Triplet-to-Track System (TTS) enhances long-horizon robotic manipulation by integrating hierarchical planning with object-centric representations, leveraging human videos to minimize dependence on robot-collected data. By representing high-level subgoals as instance-grounded triplets and employing continuous track priors for execution, TTS enables real-time monitoring and online replanning based on task progress. This approach resulted in a 74.8% average success rate across various real-world tasks, demonstrating improved reliability and generalization capabilities in uncertain environments.
TTS achieves a remarkable 74.8% success rate in long-horizon manipulation by grounding high-level goals in real-world observations, transforming how robots learn from human demonstrations.
Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but their plans are often weakly grounded in observations, weakly aligned with low-level actions, and computed without online feedback, leading to open-loop behavior and hallucinations. To address these issues, we introduce the Triplet-to-Track System (TTS), a closed-loop long-horizon imitation learning system that uses human videos to reduce reliance on robot-collected data. TTS represents high-level subgoals as instance-grounded triplets, translates them into continuous track priors for execution, and monitors task progress from observations for online replanning. Across diverse real-world long-horizon tasks, TTS achieves a 74.8\% average success rate and supports object-level and compositional generalization.