Search papers, labs, and topics across Lattice.
This paper explores the transfer of text-conditioned grasp detection models to speech inputs in humanoid robots, addressing a gap in multi-modal understanding. By employing a lightweight MLP-based projector to adapt the ALBEF model, the authors demonstrate that this approach maintains semantic discrimination and robustness while being data-efficient. The proposed Speech2Grasp framework not only surpasses traditional ASR-based methods in real-world experiments but also reduces inference latency, highlighting its practical applicability in robotic interactions.
Speech2Grasp achieves superior grasp detection from speech inputs, outperforming traditional ASR pipelines while cutting down inference latency.
Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.