Search papers, labs, and topics across Lattice.
This study investigates whether episodic memory mechanisms can mitigate the sensitivity of language models to lexical frequency in syntactic contrast tasks. By employing retrieval-augmented language models, specifically $k$-nearest-neighbor models, the researchers demonstrate that these models can significantly reduce the performance gap between high- and low-frequency lexical items in grammaticality judgments. The findings indicate that while episodic memory can enhance performance, it does not completely eliminate the frequency gap, suggesting avenues for further refinement in retrieval strategies and instance representation.
Retrieval-augmented language models can significantly narrow the lexical frequency gap in syntactic contrast sensitivity, but they still leave room for improvement.
Grammatical knowledge and how it is empirically tested are typically considered robust to the frequency of the lexical items in the expressions. However, neural network-based models of grammaticality exhibit high sensitivity to lexical frequency. We draw upon Complementary Learning Systems theory to test the hypothesis that robustness to lexical frequency can arise via a hippocampal episodic memory mechanism, which enables rapid encoding and retrieval of specific experiences and allows learners to leverage them when processing rare patterns. We use retrieval-augmented language models as an instantiation of such an episodic memory mechanism (specifically, $k$-nearest-neighbor language models that augment parametric models with explicit instance storage), and test whether this augmentation helps close the lexical frequency gap that vanilla language models exhibit in syntactic contrast tests. Using syntactic contrasts with frequency-stratified test items, we find that retrieval augmentation narrows the performance gap between high- and low-frequency items, consistent with episodic memory compensating for weak parametric representations. This benefit is consistent across different syntactic phenomena and across models pretrained on child-realistic and large-scale data. Additionally, we show that structural information is critical for effective retrieval, whereas semantic similarity alone provides little benefit. While these are promising proof-of-concept results supporting our hypothesis, the frequency gap is narrowed rather than fully closed. Based on our analyses, we propose preferential reweighting of retrieved instances, better representations and retrieval strategies for structural information, and flexible configurations of storage and retrieval as promising future directions for improving the implementation of episodic memory in language models.