Search papers, labs, and topics across Lattice.
This paper introduces HINT, a framework that enhances long-horizon robot manipulation by aligning semantic intent with evolving visual observations. By leveraging sparse semantic reasoning at manipulation-pattern transitions and maintaining commitment through multi-view grounding, HINT effectively navigates complex tasks where traditional vision-language models falter. Experimental results demonstrate significant improvements in intent understanding and task success across multiple long-horizon scenarios without compromising control latency.
HINT achieves a remarkable balance between semantic intent and visual adaptability, leading to substantial gains in robot manipulation success rates.
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual inputs and sparse language guidance. Visual correlations can then dominate semantic intent, leading actions to follow visual shortcuts rather than human goals. We present HINT (Human-INTent INcepTion), an agentic framework inspired by the human manipulation principles: semantic intent changes sparsely at manipulation-pattern transitions, whereas continuous control primarily depends on the evolving object-hand relationship. HINT invokes semantic reasoning only at pattern transitions to resolve the current subtask and target, then maintains this commitment through multi-view grounding and visual tracking. We explore two visual interfaces-image-space semantic highlighting and attention-prior injection-to communicate the tracked intent to the action policy without introducing additional trainable parameters into the foundation action model. Experiments across three long-horizon tasks and out-of-distribution variants show that HINT substantially improves intent understanding, task progress, and end-to-end success across two foundation policies while preserving low-latency control.