Search papers, labs, and topics across Lattice.
This study investigates the capacity of large language models (LLMs) to track belief states by embedding a controllable latent variable within natural text, utilizing a teacher model that subliminally steers the text along a ring-shaped Markov chain of eight directions. The findings reveal that a smaller transformer model can effectively estimate the Bayesian posterior of the planted latent variable and organize the belief states in accordance with the Markov chain's structure. This work bridges the gap between belief state representation and the geometric interpretation of concepts in LLMs, providing empirical evidence for the relationship between latent variables and feature geometry.
LLMs can not only track belief states but also geometrically organize them in a way that mirrors the underlying statistical dynamics of their latent variables.
LLMs are thought to track"belief states,"i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sarfati et al., 2026), but so far this has only been comprehensively demonstrated on toy synthetic data and in a few isolated case studies. It has also never been empirically connected to the geometry of LLM features (the concepts interpretability finds in model activations). In this work, we plant a controllable latent variable inside natural-looking text. An LLM teacher writes ordinary text while we"subliminally"steer it along one of K = 8 unrelated sparse autoencoder directions at each token, with the active directions following a ring-shaped Markov chain. A small transformer model trained on this corpus does indeed track the Bayesian posterior belief about our planted latent variable. Moreover, it also arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.