Search papers, labs, and topics across Lattice.
This study investigates the dynamics of using a multimodal large language model (MLLM) as an interviewer in semi-structured interviews, employing a voice-based system called InterviewBot. Through an empirical analysis of 15 participants, the researchers found that while the MLLM was acknowledgment-heavy, it lacked depth in probing questions and often combined multiple inquiries into single turns, leading to several data-collection breakdowns. Key insights reveal how reduced social pressure affected participant disclosure and trust, highlighting the need for improved design in AI-driven interview systems to enhance conversational quality and participant engagement.
Trust in AI interviewers hinges on perceived stakes and the subtleties of conversational grounding, not just their technical capabilities.
Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (ii) an inductive catalogue of four data-collection breakdowns (information loss, premature termination, latency, and interruption) observed in a deployed rather than simulated system; and (iii) three social dynamics from participants'reflections: disclosure calibration, where reduced social pressure coincided with shallower elaboration; institutional legitimacy, where trust tracked perceived stakes and what delegation to AI signaled about the organizer rather than conversational competence; and conversational grounding, where content-grounded paraphrase, not generic social filler, was what participants read as listening. We conclude with design implications for depth control, transparent handoffs, and non-templated listening mechanisms in human-centered interview automation.