Search papers, labs, and topics across Lattice.
This paper introduces a lightweight transformer model for predicting emotion-aware iconic gestures from text, bypassing the need for audio input during inference. The model predicts both the placement and intensity of iconic gestures based on textual and emotional cues. Results on the BEAT2 dataset show the proposed model outperforms GPT-4o in both semantic gesture placement classification and intensity regression, while maintaining computational efficiency.
Forget audio – this lightweight transformer predicts robot gestures from text and emotion better than GPT-4o.
Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that derives iconic gesture placement and intensity from text and emotion alone, requiring no audio input at inference time. The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression on the BEAT2 dataset, while remaining computationally compact and suitable for real-time deployment on embodied agents.