Search papers, labs, and topics across Lattice.
This paper introduces TacForcing, a novel streaming action-generation framework that integrates execution-time tactile feedback to enhance contact-rich manipulation tasks. By replacing traditional action experts with a streaming action expert and implementing Execution-Aware Tactile Attention (EATA), TacForcing minimizes the temporal mismatch between tactile data acquisition and action execution. The framework demonstrates significant improvements in success rates, achieving 65% in simulated tasks and 69% in real-world scenarios, outperforming existing methods.
Streaming action generation with real-time tactile feedback boosts success rates in complex manipulation tasks by over 20%.
Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework that effectively incorporates execution-time tactile feedback. Instead of employing a separate reactive controller, TacForcing replaces the standard action expert with a streaming action expert to generate actions conditioned on the evolving tactile observations acquired during execution. TacForcing also introduces Execution-Aware Tactile Attention (EATA), which restricts tactile conditioning to actions nearing execution, thereby reducing the temporal mismatch between tactile acquisition and action execution. Across six simulated UniVTAC tasks and three real-world contact-rich manipulation tasks, TacForcing achieves average success rates of 65% and 69%, respectively, outperforming strong baselines in both settings.