Search papers, labs, and topics across Lattice.
This paper addresses the challenge of maintaining persistent object states in humanoid loco-manipulation by introducing Persistent Object Tokenization (POT), which creates role-indexed 3D object records from RGB-D observations. The approach enables a closed-loop execution system where the same object tokens are used for both action generation and verification, significantly enhancing the robot's ability to manage tasks involving complex object interactions. Experimental results show that the proposed POT-VLA system outperforms existing methods, achieving a success rate of 71/80 in real-world tasks, particularly excelling in scenarios requiring sustained 3D object relations.
Persistent object tokens enable humanoid robots to achieve a 71/80 success rate in complex loco-manipulation tasks, significantly outpacing previous benchmarks.
Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts them into object tokens for a whole-body action expert. Instantiated as \emph{POT-VLA}, the same object records condition action generation and support geometric predicate checks, yielding a closed-loop execution system in which object state is both actionable and verifiable. On a Unitree G1, POT-VLA improves a matched direct GR00T-N1.7 baseline from 39/80 to 71/80 successes over eight real-world task families. In an external Being-0-aligned reference, POT-VLA achieves 44/50 successes on aligned service tasks, compared with the 37/50 success reported by the Being-0 paper. The largest gains occur on tasks requiring maintained 3D relations, suggesting that persistent object-centered state is a useful abstraction for verifiable humanoid VLA execution.