Search papers, labs, and topics across Lattice.
This paper introduces Dual-Glob, a deep supervised contrastive learning framework, to classify pitch accent patterns in Seoul Korean by capturing holistic F0 contour shapes. The method enforces structural consistency between clean and augmented views of F0 contours in a shared latent space. Experiments on a newly introduced large-scale dataset of 10,093 Accentual Phrases demonstrate that Dual-Glob significantly outperforms baseline models, achieving state-of-the-art accuracy (77.75%) and F1-score (51.54%).
Seoul Korean pitch accent classification achieves state-of-the-art results by learning F0 contour representations with deep supervised contrastive learning, despite the inherent variability in real-world speech.
The intonational structure of Seoul Korean has been defined with discrete tonal categories within the Autosegmental-Metrical model of intonational phonology. However, it is challenging to map continuous $F_0$ contours to these invariant categories due to variable $F_0$ realizations in real-world speech. Our paper proposes Dual-Glob, a deep supervised contrastive learning framework to robustly classify fine-grained pitch accent patterns in Seoul Korean. Unlike conventional local predictive models, our approach captures holistic $F_0$ contour shapes by enforcing structural consistency between clean and augmented views in a shared latent space. To this aim, we introduce the first large-scale benchmark dataset, consisting of manually annotated 10,093 Accentual Phrases in Seoul Korean. Experimental results show that our Dual-Glob significantly outperforms strong baseline models with state-of-the-art accuracy (77.75%) and F1-score (51.54%). Therefore, our work supports AM-based intonational phonology using data-driven methodology, showing that deep contrastive learning effectively captures holistic structural features of continuous $F_0$ contours.