Search papers, labs, and topics across Lattice.
This study evaluates the automatic estimation of rapport in human-robot interactions (HRI) using multimodal recordings from 62 sessions in a Japanese drugstore, addressing the challenge of applying laboratory-based evaluation methods to real-world scenarios. The research reveals that zero-shot large language models (LLMs) perform robustly in estimating rapport, while audio and visual models provide complementary insights, with the best results achieved through a fusion of Gemini 2.5 Flash and other models. Notably, the effectiveness of rapport estimation fluctuates based on interaction duration and group size, underscoring the need for context-aware evaluation methods in HRI.
Zero-shot LLMs can reliably estimate rapport in real-world human-robot interactions, outperforming traditional models in dynamic environments.
Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary information. In particular, Gemini 2.5 Flash performs strongly as a single model, and a fusion model combining Gemini (text) with HuBERT and V-JEPA performs best overall. Further analyses showed that estimation performance varied across interaction-duration and group-size conditions. These findings suggest that rapport estimation in real-world HRI requires evaluation and model design that account for contextual variability beyond that assumed in laboratory settings.