Search papers, labs, and topics across Lattice.
This study reveals that the arrangement of words in human language follows a scaling law of contextual persistence, quantified by the contextual persistence function P(d), which measures the reduction in perplexity conferred by prior context at varying distances. By employing large language models as probabilistic probes across diverse corpora, the researchers found that P(d) decays approximately as 1/d, indicating a uniform distribution of contextual influence over logarithmic timescales. This finding not only highlights the structured nature of language but also differentiates linguistic sequences from other forms of data, such as genomic sequences, underscoring the unique properties of human language processing.
Contextual influence in human language diminishes predictably with distance, following a scaling law that could redefine our understanding of linguistic structure.
Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-\alpha}$, mean $\alpha = 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.