Gachon UniversityRutgersApr 21, 2026arXiv:2604.19477

Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

Hyunjung Joo, Gyeong-Myeong Lee, GyeongTaek Lee

AI Summary

This paper introduces Dual-Glob, a deep supervised contrastive learning framework, to classify pitch accent patterns in Seoul Korean by capturing holistic F0 contour shapes. The method enforces structural consistency between clean and augmented views of F0 contours in a shared latent space. Experiments on a newly introduced large-scale dataset of 10,093 Accentual Phrases demonstrate that Dual-Glob significantly outperforms baseline models, achieving state-of-the-art accuracy (77.75%) and F1-score (51.54%).

Key Contribution

Seoul Korean pitch accent classification achieves state-of-the-art results by learning F0 contour representations with deep supervised contrastive learning, despite the inherent variability in real-world speech.

Abstract

The intonational structure of Seoul Korean has been defined with discrete tonal categories within the Autosegmental-Metrical model of intonational phonology. However, it is challenging to map continuous $F_0$ contours to these invariant categories due to variable $F_0$ realizations in real-world speech. Our paper proposes Dual-Glob, a deep supervised contrastive learning framework to robustly classify fine-grained pitch accent patterns in Seoul Korean. Unlike conventional local predictive models, our approach captures holistic $F_0$ contour shapes by enforcing structural consistency between clean and augmented views in a shared latent space. To this aim, we introduce the first large-scale benchmark dataset, consisting of manually annotated 10,093 Accentual Phrases in Seoul Korean. Experimental results show that our Dual-Glob significantly outperforms strong baseline models with state-of-the-art accuracy (77.75%) and F1-score (51.54%). Therefore, our work supports AM-based intonational phonology using data-driven methodology, showing that deep contrastive learning effectively captures holistic structural features of continuous $F_0$ contours.

Natural Language Processing Speech & Audio

Citation Metrics

Citations0

Influential citations0

References38

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

Related Papers