Search papers, labs, and topics across Lattice.
This paper critiques surprisal theory by demonstrating that it is tautological without additional constraints on language models, as any non-negative measure of processing difficulty can be mapped to an affine function of surprisal. The author highlights that the assumption linking language models to the distribution of training corpora has been challenged by recent empirical findings, which show that improved corpus fit does not necessarily enhance predictions of human processing difficulty. To resolve this tautology, the paper advocates for a rationalist approach, suggesting that language models should be derived from a theoretical understanding of comprehender behavior rather than solely from empirical data.
Surprisal theory's reliance on language models leads to tautological predictions, undermining its empirical validity in psycholinguistics.
Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.