Search papers, labs, and topics across Lattice.
This study introduces Real-TurnTurk, a multimodal dataset specifically designed for turn-taking prediction in Turkish conversations, featuring synchronized video, audio channels, and time-aligned transcriptions. By formulating turn-taking prediction as a binary classification problem, the authors employ a Genetic Algorithm to optimize interpretable decision rules based on visual, acoustic, and linguistic features. The key finding reveals that a hybrid AND-OR rule representation effectively captures the complex cue combinations that signal turn transitions, enhancing the modeling of conversational dynamics in Turkish.
Turn-taking prediction in Turkish conversations can be significantly improved using a novel multimodal dataset and a hybrid rule-based approach that captures complex interaction cues.
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for turn-ending prediction, there is a lack of naturalistic conversational corpora specifically addressing turn-taking dynamics in Turkish. This study introduces a multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions. Turn-taking prediction is formulated as a binary classification problem, and a Genetic Algorithm (GA) is employed to optimize interpretable decision rules derived from visual, acoustic, and linguistic features. A hybrid AND-OR rule representation is adopted in the proposed framework to represent the alternative cue combinations that precede a turn transition.