Search papers, labs, and topics across Lattice.
This paper introduces $\tau_0$-VLA, a hierarchical robot foundation model that enhances long-horizon manipulation by integrating world-model-guided test-time computation for high-level subtask generation. By allowing the high-level policy to allocate additional computation for complex decisions, the model improves next-subtask prediction accuracy significantly. The approach, trained on a vast dataset of over 40,000 hours of real-world data, demonstrates marked improvements in closed-loop success rates for robot manipulation tasks, particularly in challenging scenarios.
Allocating extra computation during decision-making boosts robot manipulation success rates, revealing a critical strategy for handling complex tasks.
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $\tau_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.