Search papers, labs, and topics across Lattice.
This paper introduces the BD-LSC dataset, which addresses the limitations of existing benchmarks in detecting bi-directional semantic change, particularly for words with both slang and standard meanings. By providing two complementary datasets鈥擝D-LSC for capturing sense gain, loss, and stability, and ST-WSD for fine-grained sense annotations鈥攖he authors enable a more nuanced evaluation of models in this domain. The evaluation reveals that while the few-shot GPT-4o model performs best overall, the challenge of accurately identifying rare slang senses persists, highlighting a significant area for future research.
Rare slang senses remain a significant challenge in semantic change detection, with Macro-F1 scores hovering around 0.5 across evaluated models.
Automatic semantic change detection aims to identify how word meanings shift over time, offering insights into both linguistic and societal change. Despite recent progress in computational lexical semantic change (LSC), existing benchmarks and methods struggle to capture bi-directional semantic change, particularly cases where words simultaneously gain and lose senses. This problem is especially challenging for words that have both slang and standard meanings. To address these gaps, we introduce two complementary benchmark datasets. The Bi-Directional Lexical Semantic Change (BD-LSC) dataset captures sense gain, sense loss, and stability across three time periods, enabling the study of complex semantic trajectories. The SlangTrack Word Sense Disambiguation (ST-WSD) dataset provides fine-grained, instance-level sense annotations for words combining slang and standard usages, supporting systematic benchmarking of WSD and semantic change detection models. Using these benchmarks, we systematically evaluate models across different methodological families: unsupervised clustering using contextualised embeddings, supervised machine learning, transformer-based models, and state-of-the-art large language models. Among the evaluated systems, the few-shot GPT-4o model achieved the strongest aggregate performance on Exact Sense Match (ESM) and multi-label accuracy; however, Macro-F1 scores near 0.5 across all systems show that rare slang senses remain difficult, which we identify as the central open challenge.