Search papers, labs, and topics across Lattice.
To move beyond coarse endpoint accuracy and entropy metrics in chain-of-thought evaluation, this work formulates answer-distribution trajectories that track an LLM's full predictive probability distribution over candidate answers at each reasoning step. Grounded in stochastic dynamics, this representation formalizes distinct operational phases—exploration, revision, motion, and commitment—to map how competing hypotheses evolve or collapse during reasoning. Evaluating 16 open-weight models across four reasoning benchmarks, the authors show that traces with identical answers and matching entropy curves exhibit radically different underlying trajectories that are systematically steered by training and inference configurations.
Reasoning traces with identical final answers and identical entropy curves can conceal radically different internal dynamics, from protracted multi-hypothesis competition to sudden late-stage revision.
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.