Search papers, labs, and topics across Lattice.
This study investigates the geometric representation of semantic relations in language models by analyzing how words related to a target word are positioned in semantic space. Through experiments with causal, masked, and diffusion models, the authors find that asymmetric relations are more distinctly represented than symmetric ones, with lexical information being more influential for causal models while contextual information prevails in masked and diffusion models. The findings reveal that the representation of semantic relations varies significantly, indicating limitations in how well these models learn from distributional data alone.
Asymmetric semantic relations are distinctly represented in language models, but the encoding of these relations reveals surprising limitations in their learning capabilities.
When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives. We first examine whether words standing in a particular relation to a target word~(called relata) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other. We then verify to what extent semantic spaces reflect certain well-known properties of relations, such as symmetry, asymmetry, and transitivity. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts. We conduct experiments on six semantic relations using causal, masked, and diffusion language models. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space. Asymmetric relations'properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models. Our results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.