Search papers, labs, and topics across Lattice.
The authors investigate diachronic semantic change in Sinhala from the 13th to the 20th century by coupling century-stratified static embedding alignment with contextual representations from a fine-tuned Llama-3.1-8B. To distinguish systemic semantic drift from temporary polysemic noise in data-scarce regimes, they implement a Bidirectional Semantic Impact Pruning framework utilizing Leave-One-Out influence diagnostics. Their analysis reveals that diachronic semantic drift in contextualized models is not a uniform, population-wide shift, but is disproportionately driven by a small subset of high-impact contextual instances.
Semantic drift over eight centuries is not a uniform, corpus-wide evolution, but is driven by a sparse cluster of high-impact contextual outliers that contextualized LLMs can isolate.
Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala language from the 13th to the 20th century using a multi-stage computational framework. We first align century-specific Word2Vec and FastText embeddings using Similarity Matrix Based Alignment (SMA) and Orthogonal Procrustes (OP) techniques, finding that OP alignment provides more stable neighbourhood tracking for identifying temporal similarity dips. To move beyond aggregate measures, we introduce a Bidirectional Semantic Impact Pruning approach using contextualised embeddings from a fine-tuned Llama-3.1-8B. By applying Leave-One-Out (LOO) diagnostics, we attempt to isolate influential sentences to distinguish between systemic semantic shifts and transient polysemic expansion. Our results show that semantic drift in the fine-tuned Llama-3.1-8B is not evenly distributed across all usages. Instead, a significant part of the change is driven by a smaller set of high-impact contextual instances, rather than gradual and uniform change across all occurrences. This work provides a preliminary framework for diachronic analysis in low-resource contexts, highlighting the trade-offs between model sensitivity and data availability.