Search papers, labs, and topics across Lattice.
This study examines the anaphor resolution capabilities of five large language models (LLMs) by assessing their sensitivity to cognitive factors identified in human processing, such as discourse structure and distance-based factors. By correlating model performance with human reading times and accuracy on comprehension questions, the authors reveal that certain LLMs demonstrate human-like behaviors in resolving anaphors, while others lack sensitivity to semantic interference. The key finding is that while some LLMs align with human cognitive processes, their performance is context-dependent, highlighting the limitations and conditions under which LLMs can mimic human-like resolution strategies.
Some large language models exhibit surprising human-like sensitivity to discourse factors in anaphor resolution, but falter on semantic interference.
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.