Search papers, labs, and topics across Lattice.
This study investigates the role of gestures in multimodal dialogue by analyzing how referential information is conveyed under varying partner visibility conditions. By developing models that utilize speech transcripts, skeletal gesture representations, or both, the research reveals that gestures alone can effectively predict intended referents, with multimodal fusion enhancing performance particularly in uncertain contexts. Additionally, the findings highlight the pragmatic effects of visibility on gesture use and the dynamics of interaction over repeated exchanges, contributing to both technical modeling and the understanding of human communication patterns.
Gestures alone can predict referential intent, revealing their critical role in multimodal dialogue even when speech is ambiguous.
Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.