Search papers, labs, and topics across Lattice.
This study employs topological methods to analyze the semantic structures of language models, addressing the challenges of interpreting high-dimensional embedding spaces. By comparing these structures to low-dimensional baselines such as ontologies and knowledge graphs, the authors provide a more nuanced understanding of how models conceptualize language and adapt over time. The findings reveal that multi-modal alignment tests can effectively track phrase understanding and model adaptations across different languages, enhancing the benchmarking process for language models.
Topological analysis reveals that language models can adapt their semantic structures in surprising ways, offering deeper insights into their conceptual understanding across languages.
Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important to test the model's output, but augmenting these tests by characterizing semantic structure gives more insight to how models relate abstract concepts. However, the high dimensional embedding spaces are not easy to interpret. This work demonstrates how topological methods can be used to rigorously compare these spaces to low dimensional and interpretable baselines like ontologies and curated knowledge graphs. These multi-modal alignment tests make it possible to track model adaptations and test phrase understanding across multiple languages.