Search papers, labs, and topics across Lattice.
This study investigates the role of minimal responses in counseling dialogues, revealing their importance in conveying empathy and encouraging client expression. Through a systematic cross-lingual analysis, the authors find that while minimal responses are prevalent in human-generated datasets, they are significantly underrepresented in outputs from large language models (LLMs). The research highlights that, although LLMs can generate minimal responses when prompted, they often fail to recognize the appropriate contexts for such brevity, leading to a preference for longer, content-rich replies.
LLMs struggle to generate contextually appropriate minimal responses in counseling, despite their ability to produce them on command.
In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, and encourage clients to continue expressing themselves. However, existing counseling dialogue systems and evaluation frameworks often favor explicit, content-rich replies, overlooking the interactional value of brief counselor utterances. This paper presents a systematic cross-lingual analysis of minimal responses across multiple counseling dialogue datasets. We develop a two-stage filtering method based on utterance length and content, followed by contextual verification using a large language model (LLM). Our analysis shows that minimal responses are common in human-collected datasets but substantially underrepresented in LLM-generated ones. We further evaluate current LLMs in manually curated dialogue contexts where human counselors used minimal responses. The results show that strong commercial LLMs are capable of generating minimal responses when explicitly instructed, but still struggle to determine when such responses are appropriate. Counseling-specific models trained on synthetic data perform particularly poorly, tending instead to produce longer and more information-rich responses. Moreover, LLM-based response-quality evaluation may undervalue minimal responses, even when they are interactionally appropriate.