Search papers, labs, and topics across Lattice.
This paper introduces TurnBench, a comprehensive multi-domain benchmark designed to evaluate turn-taking dynamics in spoken dialogue through a 30-hour hand-labeled corpus of dyadic conversations. The study reveals that while end-of-turn recall is consistently stable across different conversation types, the occurrence of false positives in interruption detection varies significantly, particularly in backchannel-dense interactions. Notably, human listeners initiate speaking just before the current turn ends, a nuance that current systems fail to replicate without generating excessive false positives.
TurnBench reveals that human-like turn-taking in dialogue remains elusive for AI, with systems struggling to balance accuracy and false positives in interruption detection.
Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com